Leaderboard

Human Verdicts are the ranked input. Automatic scoring is not used for ranking. Latency over 5,000 ms is flagged.

#Server VariantTranslation failuresAudited pass rateMean MER accuracyMean latencyOffending Takes
1mini_auto_bare_deg
gpt-4o-mini-transcribe · degraded · auto · bare
0(0 judged, 291 unaudited)704 ms
2mini_prod_baseline
gpt-4o-mini-transcribe · degraded · en · bare
0(0 judged, 291 unaudited)714 ms
3mini_auto_prompt_deg
gpt-4o-mini-transcribe · degraded · auto · mixed prompt
0(0 judged, 291 unaudited)717 ms
4gpt_tw_en_bare_deg
gpt-transcribe · degraded · zh-tw + en · bare
0(0 judged, 291 unaudited)726 ms
5mini_auto_bare_clean
gpt-4o-mini-transcribe · clean · auto · bare
0(0 judged, 291 unaudited)730 ms
6mini_zh_prompt_deg
gpt-4o-mini-transcribe · degraded · zh · mixed prompt
0(0 judged, 291 unaudited)753 ms
7gpt_tw_en_prompt_clean
gpt-transcribe · clean · zh-tw + en · mixed prompt
0(0 judged, 291 unaudited)754 ms
8gpt_tw_only_prompt_deg
gpt-transcribe · degraded · zh-tw · mixed prompt
0(0 judged, 291 unaudited)760 ms
9mini_auto_prompt_64k
gpt-4o-mini-transcribe · 64k · auto · mixed prompt
0(0 judged, 291 unaudited)761 ms
10mini_auto_prompt_clean
gpt-4o-mini-transcribe · clean · auto · mixed prompt
0(0 judged, 291 unaudited)768 ms
11gpt_tw_en_prompt_64k
gpt-transcribe · 64k · zh-tw + en · mixed prompt
0(0 judged, 291 unaudited)770 ms
12mini_zh_prompt_clean
gpt-4o-mini-transcribe · clean · zh · mixed prompt
0(0 judged, 291 unaudited)774 ms
13gpt_tw_en_prompt_deg
gpt-transcribe · degraded · zh-tw + en · mixed prompt
0(0 judged, 291 unaudited)774 ms

No winner selected yet.

Pick winner