nb-whisper-medium-mlx-fp16

An MLX conversion of NbAiLab/nb-whisper-medium — the National Library of Norway's Norwegian speech recognition model — for native inference on Apple Silicon.

The National Library publishes the model in PyTorch, TensorFlow, whisper.cpp and ONNX formats, and invites conversions to other formats. No MLX build of the medium model existed. This repository provides one.

  • Base model: NbAiLab/nb-whisper-medium (Apache-2.0)
  • Format: MLX, float16 — not quantized
  • Parameters: 762.3 M (947 tensors: 946 float16, 1 int64 — see note 4)
  • Weights size: approximately 1.52 GB / 1.42 GiB
  • Architecture: Whisper medium shape — 80 mel bins, 24 encoder + 24 decoder layers, d_model 1024, 16 attention heads, vocab 51,865
  • Not retrained, fine-tuned or quantized — format conversion only

On the model tree: Hugging Face's base_model_relation field only accepts finetune, adapter, merge or quantized. None of these describes a pure format conversion, so the field is deliberately omitted here and the Hub may infer a label that overstates the relationship. This model is a format conversion of the base weights — nothing more.

Why this exists

Norwegian speech recognition that runs locally matters for material you are not allowed to send to a cloud service — journalism, healthcare, legal work, footage covered by consent agreements. For Python-based workflows on Apple Silicon, MLX provides a practical way to run Whisper-class ASR locally at usable speed. (whisper.cpp is another route, and the National Library ships a GGML build for it.)

medium is the size that fits comfortably alongside other work on a 16–24 GB Mac: roughly half the weights of large, 2.2 GB peak memory, and in the measurement below it ran 1.8× faster than large at the same word error rate on read Bokmål. That result is specific to this benchmark — see the caveat under the table.

Installation

brew install ffmpeg
pip install mlx-whisper==0.4.3

ffmpeg must be on your PATH — mlx_whisper shells out to it to decode audio. With Homebrew on Apple Silicon that usually means:

export PATH="/opt/homebrew/bin:$PATH"

Usage

import mlx_whisper

result = mlx_whisper.transcribe(
    "audio.wav",
    path_or_hf_repo="viavicdev/nb-whisper-medium-mlx-fp16",
    language="no",
)
print(result["text"])

Notes

  1. No tokenizer files are needed. mlx-whisper selects its bundled multilingual tokenizer from n_vocab (51,865 → the 99-language, Whisper-v2-era multilingual vocabulary this model uses), so this repository ships only config.json and the weights.
  2. The weights file must be named weights.safetensors. mlx_whisper's loader looks for weights.safetensors, then falls back to weights.npz. A file named model.safetensors — which is what the standard conversion script writes — produces the misleading error ValueError: [load_npz] Input must be a zip file or a file-like object that can be opened with zipfile.ZipFile. This repository uses the correct name; the note is here in case you re-convert.
  3. 80 mel bins, not 128. This model uses the Whisper-v2-era front end. mlx_whisper reads that from config.json and handles it automatically, but it matters if you build your own feature extraction.
  4. Word-level timestamps use fallback alignment heads. The single non-float16 tensor is alignment_heads [192, 2]. Hugging Face checkpoints do not carry Whisper's model-specific alignment heads, so mlx_whisper falls back to its default — every head in the second half of the decoder (12 layers × 16 heads). Segment-level timestamps are unaffected; word_timestamps=True will be less precise than with calibrated heads.
  5. float16, not quantized. The checkpoint remains in float16; no low-bit quantization was applied, avoiding the additional approximation that 4-bit quantization introduces. If you need a smaller footprint, quantize it yourself.

Measured behaviour

A documented measurement on a public benchmark, on one machine. Not a leaderboard submission.

Material: google/fleurs nb_no, validation split, CC-BY-4.0. Every sentence in the split appears two or three times with different readers; the set used here is one utterance per unique sentence id (the first row per id in dev.tsv), which gives 111 utterances / 1,415.0 seconds of read Bokmål. Deduplicating matters: without it some sentences are weighted two or three times.

Environment: Apple M4 Pro, 24 GB unified memory · macOS 15.7.2 · Python 3.12.12 · mlx 0.32.0 · mlx-whisper 0.4.3 · ffmpeg 8.1.2

Scoring: word error rate against the dataset's transcription column, with the same normalisation applied to reference and hypothesis — lowercased, punctuation stripped, whitespace collapsed. Numbers are not normalised, so 15 versus femten counts as an error.

nb-whisper-medium-mlx-fp16 nb-whisper-large, MLX fp16
WER 7.00 % 7.00 %
Errors / reference words 165 / 2,358 165 / 2,358
Utterances with zero errors 42 / 111 44 / 111
Total inference time 72.8 s 131.2 s
Speed 19.5× real time 10.8× real time
Cold model load 4.7 s 9.6 s
Peak memory 2.2 GB not recorded

Model load is excluded from the timings; ffmpeg audio decoding is included. Both models were run on the same audio, on the same machine, on the same day, through the same script.

The identical WER is a coincidence of aggregation, not evidence that the two models behave the same. They produced different transcriptions for 45 of the 111 utterances and different error counts for 41 of them. Medium was better on 21 utterances, large on 20, and they tied on 70. Each truncated exactly one utterance. The totals happened to land on the same number.

What the measurement does support: on read Bokmål, medium costs nothing measurable in accuracy while running 1.8× faster in half the memory. On harder material — spontaneous speech, noise, overlapping speakers, dialect — large may well pull ahead. That was not tested here.

What the errors look like

Of the 165 counted errors, a substantial share are orthographic and formatting differences rather than misheard speech:

  • Number and abbreviation spelling — 15 of the single-word substitutions involve digits: 3 → tre, 2 → to, nr → nummer, 100 m → 100 meter, på grunn av → pga.
  • Inflection — vitenskapelig forskning → vitenskapelige forskningen, statlige → statlig.
  • Proper nouns — Aucklands → Auklands, Komorene → Comorene, Krezel → Kresel.

Genuine content errors do occur (ufint → en ufin, rumenere → rumerne), mostly on foreign names and dense subordinate clauses.

An example transcribed with zero errors (30 words, 16.2 s):

den nye befolkningen vil trenge ulike funksjoner eller tilpasninger enn det de trengte før for å være en sterk konkurrent siden dette nye miljøet har ulike ressurser og ulike konkurrenter

Behaviour on silence

Thirty seconds of digital silence produced ! and nothing else — no fabricated sentence. This is one test, not a guarantee; see the limitation below.

Limitations

  • Whisper-family models hallucinate on silence, music and ambience. The silence test above came out clean, but the failure mode is real for the whole family and is documented for the base model. On music, room tone and long pauses, expect invented sentences, repetition loops and spurious credit lines («Teksting av …»), and filter for them.
  • Bokmål only, so far. The base model covers Bokmål and Nynorsk; this conversion was measured on Bokmål. No Nynorsk test was run.
  • Read speech only. FLEURS is read speech. Spontaneous, overlapping, accented or noisy speech is harder and was not evaluated here.
  • It can stop early. One utterance in 111 was truncated — the model returned 9 words where the reference has 16, and simply ended. large did the same on a different utterance. If completeness matters, check output length against audio duration rather than assuming a returned transcript is a finished one.
  • float16 only. No quantized variant is provided here.
  • No speaker diarization. It transcribes what is said, not who says it.
  • One machine, one benchmark. All figures above come from a single M4 Pro and one public read-speech set.

Companion model

License and attribution

This conversion inherits the Apache-2.0 license of the base model.

All credit for the original model training and weights belongs to the NB-Whisper team at the National Library of Norway — the NoSTraM project, led by Per Egil Kummervold. See NbAiLab/nb-whisper-medium.

This repository provides the MLX format conversion, packaging, Apple Silicon compatibility testing and usage documentation. The model was not retrained, fine-tuned or quantized.

Test material in the measurements above is from google/fleurs (Conneau et al.), licensed CC-BY-4.0.

When citing the model, cite the original authors.

Downloads last month
13
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ensamble-as/nb-whisper-medium-mlx-fp16

Finetuned
(7)
this model