fast-ukrainian-asr-600m β Ukrainian speech recognition, 600M, CTC
A Ukrainian recogniser built by adapting GigaAM's 600M multilingual SSL encoder, then adapting it to telephone audio. It is the larger sibling of fast-ukrainian-asr (220M) and is more accurate than it on almost every measurement below.
It is the most accurate model we tested on read and telephone Ukrainian, and it is not the most accurate on spontaneous speech. Whisper large-v3 is still clearly better there. The table says so; please read it before deploying.
Results
750 clips (250 per set), Ukrainian-tagged rows only. Every system is scored
through the same normalizer (GigaAM's normalize_raw_text) because Whisper
emits punctuation and capitals that a 38-character CTC head never produces, so
raw-string WER would measure formatting rather than recognition. The telephony
condition is a G.711 path β band-limit, 16kβ8kβ16k, mu-law β generated once
and shared byte-identically by every system, rather than re-jittered per model.
| model | cv10 clean/tel | rs-test clean/tel | test-y clean/tel | 1500 utts |
|---|---|---|---|---|
| this model (600M) | 6.13 / 7.21 | 26.35 / 27.41 | 26.59 / 28.48 | 1381 s (CPU) |
| fast-ukrainian-asr (220M) | 9.64 / 10.59 | 28.03 / 28.65 | 26.28 / 28.67 | 576 s (CPU) |
| whisper-large-v3-turbo-uk fine-tune | 10.91 / 13.59 | 22.96 / 23.56 | 23.47 / 28.24 | 772 s (GPU) |
| whisper-large-v3 | 13.85 / 17.42 | 16.96 / 19.27 | 20.90 / 28.06 | 1051 s (GPU) |
| Qwen3-ASR | 80.73 / 83.15 | 78.77 / 77.22 | 88.39 / 84.23 | 225 s (GPU) |
cv10 is where it wins, and cv10 is Common Voice. 6.13% clean is less than
half Whisper large-v3's 13.85%. Common Voice Ukrainian is also part of this
model's training data β as a corpus, not as these utterances: sentence overlap
with the evaluation set was measured at 1 sentence in 3,204, and the Common Voice
validated and test splits were excluded from training precisely because they
share 1,326 and 396 sentences with it. So this is genuine domain match, not
leakage β but it is domain match, and a model trained on read speech scoring well
on read speech is the least surprising row in the table.
rs-test and test-y are the honest columns, because nothing related to them was trained on. There Whisper large-v3 leads by 6β10 points on clean audio. This model has seen broadcast and read speech; conversational audio is its weak case.
The telephony column is why it exists. Clean β phone costs this model +1.08 WER on cv10. It costs Whisper large-v3 +3.57. On test-y telephony the gap to Whisper closes to 0.42 points, from 5.69 on clean. For a phone product that robustness matters more than a better clean number.
Against the 220M it wins five of six cells, by 3.4β3.5 points on cv10 and 1.2β1.7 on rs-test. It loses test-y clean by 0.31 β inside the noise of 93 utterances, but reported rather than hidden. The cost is speed: 2.4Γ the CPU time. We did not benchmark it on GPU, so no GPU figure is quoted here.
Qwen3-ASR does not support Ukrainian at all β its 30-language list has no entry for it, so it transcribes Ukrainian speech into Russian orthography. The row exists so nobody repeats the experiment.
Scope: Ukrainian only
Deliberately monolingual. On Russian utterances it scores ~89% WER and always will: the vocabulary is 38 Ukrainian characters with no Russian-only letters. This matters when reading benchmarks β much Ukrainian evaluation data is quietly mixed. Of the sets above, rs-test is 17% Russian and test-y is 49% Russian, so their whole-set WERs understate any Ukrainian-only model badly. Every number here is Ukrainian-tagged rows only. If your audio is mixed, use a multilingual model.
Training
- Backbone: GigaAM
multilingual_large_ssl(585.3M params), an SSL encoder pretrained on 2M hours across 70+ languages, none of them Ukrainian. - Stage 1: 350 h β Yehor/broadcast-speech-uk (309 h) + Common Voice 22 Ukrainian (41 h) + tg-voices-uk, 181,479 clips. 8 epochs, effective batch 16, charwise CTC head, 38 classes.
- Stage 2: telephony adaptation, 4 epochs at lr 3e-5, 60% of training clips degraded through the phone path and validation 100% degraded, so checkpoint selection ranks on the condition production serves.
Common Voice's train and other splits only (see above). 146 rows containing
Latin/Cyrillic homoglyphs β a Latin i typed where Cyrillic Ρ belongs β were
repaired rather than dropped, since the vocabulary contains no Latin at all and
would otherwise have been trained to emit a character that is then scored wrong.
Usage
import gigaam
model = gigaam.load_model("fast-ukrainian-asr-600m.ckpt", device="cuda") # or "cpu"
print(model.transcribe("call.wav"))
longform.py handles files longer than ~20 s by energy-based segmentation.
server.py exposes an OpenAI-compatible /v1/audio/transcriptions endpoint.
Audio must be 16 kHz mono; 8 kHz telephone audio should be upsampled, not fed at
its native rate.
Limitations
- Spontaneous/conversational speech is its weakest case (see rs-test, test-y).
- Strongest on Common-Voice-like read speech, which is partly a domain match.
- No punctuation or capitalisation β CTC over 38 characters.
- No Russian, by design.
- Numbers are transcribed as words, not digits.
- 2.4Γ the CPU cost of the 220M; if latency matters more than accuracy, use that.
- Downloads last month
- 43