fast-ukrainian-asr-600m β€” Ukrainian speech recognition, 600M, CTC

A Ukrainian recogniser built by adapting GigaAM's 600M multilingual SSL encoder, then adapting it to telephone audio. It is the larger sibling of fast-ukrainian-asr (220M) and is more accurate than it on almost every measurement below.

It is the most accurate model we tested on read and telephone Ukrainian, and it is not the most accurate on spontaneous speech. Whisper large-v3 is still clearly better there. The table says so; please read it before deploying.

Results

750 clips (250 per set), Ukrainian-tagged rows only. Every system is scored through the same normalizer (GigaAM's normalize_raw_text) because Whisper emits punctuation and capitals that a 38-character CTC head never produces, so raw-string WER would measure formatting rather than recognition. The telephony condition is a G.711 path — band-limit, 16k→8k→16k, mu-law — generated once and shared byte-identically by every system, rather than re-jittered per model.

model cv10 clean/tel rs-test clean/tel test-y clean/tel 1500 utts
this model (600M) 6.13 / 7.21 26.35 / 27.41 26.59 / 28.48 1381 s (CPU)
fast-ukrainian-asr (220M) 9.64 / 10.59 28.03 / 28.65 26.28 / 28.67 576 s (CPU)
whisper-large-v3-turbo-uk fine-tune 10.91 / 13.59 22.96 / 23.56 23.47 / 28.24 772 s (GPU)
whisper-large-v3 13.85 / 17.42 16.96 / 19.27 20.90 / 28.06 1051 s (GPU)
Qwen3-ASR 80.73 / 83.15 78.77 / 77.22 88.39 / 84.23 225 s (GPU)

cv10 is where it wins, and cv10 is Common Voice. 6.13% clean is less than half Whisper large-v3's 13.85%. Common Voice Ukrainian is also part of this model's training data β€” as a corpus, not as these utterances: sentence overlap with the evaluation set was measured at 1 sentence in 3,204, and the Common Voice validated and test splits were excluded from training precisely because they share 1,326 and 396 sentences with it. So this is genuine domain match, not leakage β€” but it is domain match, and a model trained on read speech scoring well on read speech is the least surprising row in the table.

rs-test and test-y are the honest columns, because nothing related to them was trained on. There Whisper large-v3 leads by 6–10 points on clean audio. This model has seen broadcast and read speech; conversational audio is its weak case.

The telephony column is why it exists. Clean β†’ phone costs this model +1.08 WER on cv10. It costs Whisper large-v3 +3.57. On test-y telephony the gap to Whisper closes to 0.42 points, from 5.69 on clean. For a phone product that robustness matters more than a better clean number.

Against the 220M it wins five of six cells, by 3.4–3.5 points on cv10 and 1.2–1.7 on rs-test. It loses test-y clean by 0.31 β€” inside the noise of 93 utterances, but reported rather than hidden. The cost is speed: 2.4Γ— the CPU time. We did not benchmark it on GPU, so no GPU figure is quoted here.

Qwen3-ASR does not support Ukrainian at all β€” its 30-language list has no entry for it, so it transcribes Ukrainian speech into Russian orthography. The row exists so nobody repeats the experiment.

Scope: Ukrainian only

Deliberately monolingual. On Russian utterances it scores ~89% WER and always will: the vocabulary is 38 Ukrainian characters with no Russian-only letters. This matters when reading benchmarks β€” much Ukrainian evaluation data is quietly mixed. Of the sets above, rs-test is 17% Russian and test-y is 49% Russian, so their whole-set WERs understate any Ukrainian-only model badly. Every number here is Ukrainian-tagged rows only. If your audio is mixed, use a multilingual model.

Training

  • Backbone: GigaAM multilingual_large_ssl (585.3M params), an SSL encoder pretrained on 2M hours across 70+ languages, none of them Ukrainian.
  • Stage 1: 350 h β€” Yehor/broadcast-speech-uk (309 h) + Common Voice 22 Ukrainian (41 h) + tg-voices-uk, 181,479 clips. 8 epochs, effective batch 16, charwise CTC head, 38 classes.
  • Stage 2: telephony adaptation, 4 epochs at lr 3e-5, 60% of training clips degraded through the phone path and validation 100% degraded, so checkpoint selection ranks on the condition production serves.

Common Voice's train and other splits only (see above). 146 rows containing Latin/Cyrillic homoglyphs β€” a Latin i typed where Cyrillic Ρ– belongs β€” were repaired rather than dropped, since the vocabulary contains no Latin at all and would otherwise have been trained to emit a character that is then scored wrong.

Usage

import gigaam
model = gigaam.load_model("fast-ukrainian-asr-600m.ckpt", device="cuda")  # or "cpu"
print(model.transcribe("call.wav"))

longform.py handles files longer than ~20 s by energy-based segmentation. server.py exposes an OpenAI-compatible /v1/audio/transcriptions endpoint. Audio must be 16 kHz mono; 8 kHz telephone audio should be upsampled, not fed at its native rate.

Limitations

  • Spontaneous/conversational speech is its weakest case (see rs-test, test-y).
  • Strongest on Common-Voice-like read speech, which is partly a domain match.
  • No punctuation or capitalisation β€” CTC over 38 characters.
  • No Russian, by design.
  • Numbers are transcribed as words, not digits.
  • 2.4Γ— the CPU cost of the 220M; if latency matters more than accuracy, use that.
Downloads last month
43
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Datasets used to train asfberlin/fast-ukrainian-asr-600m