Khmer (ខ្មែរ) review centers on this evidence: inspect Khmer dependent vowels, consonant series, word segmentation, names, particles, and formal versus conversational vocabulary. CLDR maximizes bare km to Khmr (Khmer script) as default web-locale metadata. It is not evidence about the recording or proof that Khmer has only one writing system. Confirm the actual orthography and script with the project owner before normalizing any text. Orthographic targets include word segmentation, dependent vowels, names, number forms, and sentence boundaries. Project decision: which Khmer register, word-boundary, and number conventions the export follows. Before batch editing km, a reviewer must record which Khmer register, word-boundary, and number conventions the export follows and attach a source timestamp to disputed text.
Where this workflow fits
- Khmer interviews and oral histories
- Khmer lectures and training recordings
- Source-linked Khmer research review
Khmer orthography and speech checkpoints
inspect Khmer dependent vowels, consonant series, word segmentation, names, particles, and formal versus conversational vocabulary. Editorial decision: which Khmer register, word-boundary, and number conventions the export follows. Visual inspection should cover record the project decision about which Khmer register, word-boundary, and number conventions the export follows.
Script choice and source-linked review for ខ្មែរ
CLDR maximizes bare km to Khmr (Khmer script) as default web-locale metadata. It is not evidence about the recording or proof that Khmer has only one writing system. Confirm the actual orthography and script with the project owner before normalizing any text. For field interviews where word segmentation is editorially important, inspect inspect Khmer dependent vowels, consonant series, word segmentation, names, particles, and formal versus conversational vocabulary. Visually inspect word segmentation, dependent vowels, names, number forms, and sentence boundaries; record the project decision about which Khmer register, word-boundary, and number conventions the export follows. Replay a complete clause around every uncertain name, quantity, interruption, or ending.
Frozen deployment boundary
This page is published only because its language is present in capability snapshot m10-2026-07-20.1. It does not extend that catalog or claim a measured accuracy score.
- The km hint maps to the reviewed km provider root and preserves km as the user's transcript language tag.
- CLDR Khmr is a likely/default locale hint, not an audio property; capability flags remain per language and model while editors confirm the project's real script and orthography.
Reviewed provider capability
assemblyai · universal-2
root km · tiers fast, standard, precision
- Automatic detection
- Reviewed no
- Diarization
- Reviewed no
- Word timestamps
- Reviewed yes
Questions from the workflow
What must a Khmer editor decide before correcting the transcript?
written boundaries, dependent signs, names, and formal versus conversational choices need sentence-level playback. For Khmer, inspect Khmer dependent vowels, consonant series, word segmentation, names, particles, and formal versus conversational vocabulary. Project choice: which Khmer register, word-boundary, and number conventions the export follows. Keep unresolved forms beside their source timestamps.
Reviewed sources
- IANA Language Subtag Registry · checked 2026-07-20
- AssemblyAI model and supported-language documentation · checked 2026-07-20
- AssemblyAI pre-recorded audio supported languages · checked 2026-07-20
- Unicode CLDR likely-subtags language and script data · checked 2026-07-20
- Khmer language and orthography editorial reference · checked 2026-07-20
- World Atlas of Language Structures Online · checked 2026-07-20