Accuracy methodology
A benchmark should be reproducible before it is impressive.
This page documents the benchmark protocol before it publishes any scores. There is no ranking, winner, or placeholder accuracy claim.
noindex · no result claims
A result table will appear only after the dataset rights, normalization rules, model versions, run date, raw outputs, and review record are all available. Until then this page remains noindex.
- 01
Freeze an auditable corpus
Use consented or correctly licensed recordings across the reviewed fixture languages, with a frozen split and human-produced reference transcripts.
- 02
Define WER before running it
Publish the exact normalization version for case, punctuation, numerals, fillers, and script-specific tokenization before calculating word error rate.
- 03
Separate words from speakers
Score speaker attribution separately from lexical accuracy. Report overlap handling, collar, unmapped speakers, and the diarization metric used.
- 04
Preserve the run context
Record provider, model, capability snapshot, language mode, diarization setting, run date, duration, failures, latency, and billed cost for every run.
Publication gate
- 01Report per-language distributions and sample counts, not only one pooled average.
- 02Keep failures in the denominator and explain exclusions before looking at results.
- 03Publish confidence intervals or repeated-run variance where provider behavior is non-deterministic.
- 04Link each public conclusion to a versioned artifact that can be reproduced.
No benchmark result has been published.
A result table will appear only after the dataset rights, normalization rules, model versions, run date, raw outputs, and review record are all available. Until then this page remains noindex.