Instructions to use ensamble-as/nb-whisper-medium-mlx-fp16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ensamble-as/nb-whisper-medium-mlx-fp16 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir nb-whisper-medium-mlx-fp16 ensamble-as/nb-whisper-medium-mlx-fp16
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
nb-whisper-medium-mlx-fp16
An MLX conversion of NbAiLab/nb-whisper-medium — the National Library of Norway's Norwegian speech recognition model — for native inference on Apple Silicon.
The National Library publishes the model in PyTorch, TensorFlow, whisper.cpp
and ONNX formats, and invites conversions to other formats. No MLX build of the
medium model existed. This repository provides one.
- Base model: NbAiLab/nb-whisper-medium (Apache-2.0)
- Format: MLX, float16 — not quantized
- Parameters: 762.3 M (947 tensors: 946 float16, 1 int64 — see note 4)
- Weights size: approximately 1.52 GB / 1.42 GiB
- Architecture: Whisper medium shape — 80 mel bins, 24 encoder + 24 decoder layers, d_model 1024, 16 attention heads, vocab 51,865
- Not retrained, fine-tuned or quantized — format conversion only
On the model tree: Hugging Face's
base_model_relationfield only acceptsfinetune,adapter,mergeorquantized. None of these describes a pure format conversion, so the field is deliberately omitted here and the Hub may infer a label that overstates the relationship. This model is a format conversion of the base weights — nothing more.
Why this exists
Norwegian speech recognition that runs locally matters for material you are
not allowed to send to a cloud service — journalism, healthcare, legal work,
footage covered by consent agreements. For Python-based workflows on Apple
Silicon, MLX provides a practical way to run Whisper-class ASR locally at usable
speed. (whisper.cpp is another route, and the National Library ships a GGML
build for it.)
medium is the size that fits comfortably alongside other work on a 16–24 GB
Mac: roughly half the weights of large, 2.2 GB peak memory, and in the
measurement below it ran 1.8× faster than large at the same word error
rate on read Bokmål. That result is specific to this benchmark — see the
caveat under the table.
Installation
brew install ffmpeg
pip install mlx-whisper==0.4.3
ffmpeg must be on your PATH — mlx_whisper shells out to it to decode
audio. With Homebrew on Apple Silicon that usually means:
export PATH="/opt/homebrew/bin:$PATH"
Usage
import mlx_whisper
result = mlx_whisper.transcribe(
"audio.wav",
path_or_hf_repo="viavicdev/nb-whisper-medium-mlx-fp16",
language="no",
)
print(result["text"])
Notes
- No tokenizer files are needed.
mlx-whisperselects its bundled multilingual tokenizer fromn_vocab(51,865 → the 99-language, Whisper-v2-era multilingual vocabulary this model uses), so this repository ships onlyconfig.jsonand the weights. - The weights file must be named
weights.safetensors.mlx_whisper's loader looks forweights.safetensors, then falls back toweights.npz. A file namedmodel.safetensors— which is what the standard conversion script writes — produces the misleading errorValueError: [load_npz] Input must be a zip file or a file-like object that can be opened with zipfile.ZipFile. This repository uses the correct name; the note is here in case you re-convert. - 80 mel bins, not 128. This model uses the Whisper-v2-era front end.
mlx_whisperreads that fromconfig.jsonand handles it automatically, but it matters if you build your own feature extraction. - Word-level timestamps use fallback alignment heads. The single non-float16
tensor is
alignment_heads[192, 2]. Hugging Face checkpoints do not carry Whisper's model-specific alignment heads, somlx_whisperfalls back to its default — every head in the second half of the decoder (12 layers × 16 heads). Segment-level timestamps are unaffected;word_timestamps=Truewill be less precise than with calibrated heads. - float16, not quantized. The checkpoint remains in float16; no low-bit quantization was applied, avoiding the additional approximation that 4-bit quantization introduces. If you need a smaller footprint, quantize it yourself.
Measured behaviour
A documented measurement on a public benchmark, on one machine. Not a leaderboard submission.
Material: google/fleurs
nb_no, validation split, CC-BY-4.0. Every sentence in the split appears two or
three times with different readers; the set used here is one utterance per
unique sentence id (the first row per id in dev.tsv), which gives 111
utterances / 1,415.0 seconds of read Bokmål. Deduplicating matters: without it
some sentences are weighted two or three times.
Environment: Apple M4 Pro, 24 GB unified memory · macOS 15.7.2 · Python
3.12.12 · mlx 0.32.0 · mlx-whisper 0.4.3 · ffmpeg 8.1.2
Scoring: word error rate against the dataset's transcription column, with
the same normalisation applied to reference and hypothesis — lowercased,
punctuation stripped, whitespace collapsed. Numbers are not normalised, so
15 versus femten counts as an error.
| nb-whisper-medium-mlx-fp16 | nb-whisper-large, MLX fp16 | |
|---|---|---|
| WER | 7.00 % | 7.00 % |
| Errors / reference words | 165 / 2,358 | 165 / 2,358 |
| Utterances with zero errors | 42 / 111 | 44 / 111 |
| Total inference time | 72.8 s | 131.2 s |
| Speed | 19.5× real time | 10.8× real time |
| Cold model load | 4.7 s | 9.6 s |
| Peak memory | 2.2 GB | not recorded |
Model load is excluded from the timings; ffmpeg audio decoding is included. Both models were run on the same audio, on the same machine, on the same day, through the same script.
The identical WER is a coincidence of aggregation, not evidence that the two models behave the same. They produced different transcriptions for 45 of the 111 utterances and different error counts for 41 of them. Medium was better on 21 utterances, large on 20, and they tied on 70. Each truncated exactly one utterance. The totals happened to land on the same number.
What the measurement does support: on read Bokmål,
mediumcosts nothing measurable in accuracy while running 1.8× faster in half the memory. On harder material — spontaneous speech, noise, overlapping speakers, dialect —largemay well pull ahead. That was not tested here.
What the errors look like
Of the 165 counted errors, a substantial share are orthographic and formatting differences rather than misheard speech:
- Number and abbreviation spelling — 15 of the single-word substitutions
involve digits:
3→tre,2→to,nr→nummer,100 m→100 meter,på grunn av→pga. - Inflection —
vitenskapelig forskning→vitenskapelige forskningen,statlige→statlig. - Proper nouns —
Aucklands→Auklands,Komorene→Comorene,Krezel→Kresel.
Genuine content errors do occur (ufint → en ufin, rumenere → rumerne),
mostly on foreign names and dense subordinate clauses.
An example transcribed with zero errors (30 words, 16.2 s):
den nye befolkningen vil trenge ulike funksjoner eller tilpasninger enn det de trengte før for å være en sterk konkurrent siden dette nye miljøet har ulike ressurser og ulike konkurrenter
Behaviour on silence
Thirty seconds of digital silence produced ! and nothing else — no fabricated
sentence. This is one test, not a guarantee; see the limitation below.
Limitations
- Whisper-family models hallucinate on silence, music and ambience. The silence test above came out clean, but the failure mode is real for the whole family and is documented for the base model. On music, room tone and long pauses, expect invented sentences, repetition loops and spurious credit lines («Teksting av …»), and filter for them.
- Bokmål only, so far. The base model covers Bokmål and Nynorsk; this conversion was measured on Bokmål. No Nynorsk test was run.
- Read speech only. FLEURS is read speech. Spontaneous, overlapping, accented or noisy speech is harder and was not evaluated here.
- It can stop early. One utterance in 111 was truncated — the model returned
9 words where the reference has 16, and simply ended.
largedid the same on a different utterance. If completeness matters, check output length against audio duration rather than assuming a returned transcript is a finished one. - float16 only. No quantized variant is provided here.
- No speaker diarization. It transcribes what is said, not who says it.
- One machine, one benchmark. All figures above come from a single M4 Pro and one public read-speech set.
Companion model
viavicdev/nb-whisper-large-mlx-fp16— thelargemodel, same conversion procedure, same float16 policy, measured on the same set.
License and attribution
This conversion inherits the Apache-2.0 license of the base model.
All credit for the original model training and weights belongs to the NB-Whisper team at the National Library of Norway — the NoSTraM project, led by Per Egil Kummervold. See NbAiLab/nb-whisper-medium.
This repository provides the MLX format conversion, packaging, Apple Silicon compatibility testing and usage documentation. The model was not retrained, fine-tuned or quantized.
Test material in the measurements above is from
google/fleurs (Conneau et al.),
licensed CC-BY-4.0.
When citing the model, cite the original authors.
- Downloads last month
- 13
Quantized