--- license: apache-2.0 language: - 'no' - nb - nn base_model: ltg/norbert3-xs pipeline_tag: token-classification library_name: onnx datasets: - NbAiLab/NCC tags: - punctuation - truecasing - token-classification - onnx - norwegian - speech-recognition-postprocessing - kviskr --- # Kviskr tegnsetting (Norwegian punctuation and casing) Adds punctuation and capitalization to raw Norwegian speech-to-text output, without ever changing a word. Built for [Kviskr](https://kviskr.no), a macOS dictation app, where it runs on-device after NB-Whisper. - **Base:** [`ltg/norbert3-xs`](https://huggingface.co/ltg/norbert3-xs), 15M parameters - **Size:** 18.8 MB (INT8 ONNX), 20.1 MB including tokenizer - **Languages:** Norwegian Bokmål and Nynorsk - **Replaces:** [`skolfus/kviskr-punctuation-no`](https://huggingface.co/skolfus/kviskr-punctuation-no) (XLM-R base, 275 MiB) ## How it works The model predicts two labels for each whitespace-separated word: the punctuation mark that follows it, and its casing. The output is rebuilt from the original words and the original whitespace, so the text is preserved character for character. | | | |---|---| | Input | `input_ids` [B, T] int64, `attention_mask` [B, T] int64 | | Output | `tegn` [B, T, 7], `stor` [B, T, 5] (logits) | | Tokenizer | `ltg/norbert3-xs` (WordPiece, 50,000) | | Punctuation classes | none · `,` · `.` · `?` · `!` · `:` · `;` (`!` disabled in calibration) | | Casing classes | lower · Upper · UPPER · SEGMENT (`TV-aksjonen`) · BOTH (`Nord-Norge`) | ## Usage 1. Split the input on whitespace, but keep the whitespace itself. Strip trailing `.,!?;:` from each word and remember what you removed. 2. Feed the words lowercased, with `is_split_into_words=True`. 3. Read the logits on the first subword of each word. 4. Rebuild the text from the original words and whitespace. If a word already had a trailing mark, keep it and ignore the prediction for that word. 5. Change casing only when the change is reversible: every changed character must be the original's own upper or lower form, and the reverse operation must give back the original. Never touch a word that already has an internal capital (`iPhone`, `NB-Whisper`, `USA`). Keep a guard that compares the word sequence before and after and falls back to the input if anything differs. With the rules above, it should never trigger. **Word preservation, measured:** 0 violations across 3,833 texts (375 ASR clips with Ordrett input, 375 with Standard input, 3,000 text chunks and 83 dictations), under four definitions: strict, symmetric, whitespace-exact, and an independent character-level check. ## Evaluation (13 Sep 2026) All candidates receive the same ASR text; the difference is the punctuation layer alone. 95% paired bootstrap. | Benchmark | n | Previous model | This model | Paired difference | |---|---|---|---|---| | NB benchmark, NB-Whisper large verbatim | 327 clips | 58.2 | **63.7** | **+5.4** [1.2, 9.5] | | Spontaneous speech (Storting) | 60 | 47.3 | **58.1** | **+10.8** [3.1, 18.6] | | Read news prose (FLEURS) | 150 | 72.7 | 71.7 | −0.9 [−6.0, 4.3] | | Held-out Norwegian text (NCC), leak-free subset | 2,484 chunks | 61.4 | **75.5** | **+14.1** [12.8, 15.5] | | Dictation text | 83 | 77.6 | **91.1** | **+13.5** [4.7, 23.3] | Casing F1 on the NB benchmark: 77.5 → **81.1**. Full-stop F1 on the text benchmark: 88.5 → **96.1**. WER change: +0.0 [0.0, 0.1]. 504 of the 3,000 text chunks share at least one 10-gram with the training data, so the leak-free subset above is the number to cite. ## Training Fine-tuned from `ltg/norbert3-xs` on 28.2M words of Norwegian text with human punctuation from [NCC](https://huggingface.co/datasets/NbAiLab/NCC): Storting documents (13.8M), books (8.9M), Målfrid (3.9M) and Wikipedia (1.5M). 10,842 steps, batch 64, AdamW, cosine schedule. Decision thresholds calibrated per class on a held-out validation set. Norwegian speech corpora cannot serve as ground truth for punctuation: NPSC has 0.0 commas per 1,000 words, NST 17.3, and NB Tale lacks punctuation in 48 of 49 clips in the National Library's evaluation set. ## Limitations - **The two heads are independent** and can disagree: 6.4% of predicted commas on ASR text are followed by a capital letter (reference: 1.6%). - **Slightly more commas than humans on real ASR text:** 49.4 per 1,000 words against 43.1 in the reference. - **Internal capitals cannot be recovered** from lowercase input: `iphone` becomes `Iphone`. - **One- and two-word dictations get the wrong end mark** (`hei` → `Hei?`). Training chunks were 8–70 words, and 15% of real dictations are three words or fewer. ## License Apache 2.0, the same as `ltg/norbert3-xs` and the NCC sources it was trained on.