Instructions to use somniusx/recmeets-laya-words with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use somniusx/recmeets-laya-words with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Run 7 is the model RecMeets downloads since v0.25.0 (2026-10-01, Lef's choice). It's better at Greek from other speakers and microphones than run 4, and a little worse at English; see "Run 7" below. Run 4 stays in this repo's history (commit
cf3c7ee).
RecMeets' Laya: is this word a speech-recognition mistake?
A fine-tune of Laya (laya-multilingual, Apache 2.0) for RecMeets' Check words: given a word in its sentence as a speech recognizer wrote it, how likely is it a mistake? Greek and English.
Input (Laya's noul question, one per word): the question in laya.json, and this state:
Sentence: <about ten words either side, the word marked β¦ β§>
Word: <the word>
Engine confidence: <0.00β1.00, or unknown>
In the <Greek|English> dictionary: <yes|no|can't tell>
Output: the probability of "true" (a mistake), after the temperature in laya.json.
Data
Real mistakes of the engine RecMeets uses (whisper.cpp, large-v3-turbo, as RecMeets configures it): FLEURS Greek and English (Google, CC BY 4.0) transcribed in 10-minute files, each word labelled by aligning the transcript with FLEURS' reference text. Train: 71,906 Greek and 54,845 English words (9,880 and 2,547 mistakes), balanced to 35% mistakes. Test: FLEURS' test split (other speakers). The dictionary answer comes from LibreOffice's el_GR and en_US hunspell dictionaries.
Results (AUROC on FLEURS test)
| Greek, all words | Greek, unsure words | English, unsure words | |
|---|---|---|---|
| the engine's confidence alone | 0.82 | 0.63 | 0.74 |
| the dictionary alone | 0.85 | 0.70 | |
| stock laya-multilingual (same question) | 0.41 | ||
| this model | 0.90 | 0.91 | 0.86 |
| TypeSafe Jev (cloud), the same input | 0.93 | 0.89 |
"Unsure words": the engine's confidence under 0.5, four letters or more (what Check words lists).
Files
model.onnx(343 MB): exported with Laya'sexport_onnx.py; matrix products in ONNX Runtime's 8-bit block format (MatMulNBits), with 8-bit arithmetic (accuracy_level4: about 110 ms a sequence on 4 processor threads, against 140 ms without); the embedding table in 8 bits; the same AUROC as the 32-bit model (0.905 Greek, 0.860 English). Inputsinput_ids,attention_mask,marker_pos,marker_mask,qtype(2), outputlogits. One sequence at a time (the export keeps batch size 1 inside its rotary code).tokenizer.json: mmBERT's, unchanged.laya.json: the question, the token budgets, special tokens and the fitted temperature.
How it was made: scripts/laya/ and docs/research/laya-finetune.md in the RecMeets repository. By Lefteris Iliadis (Lefteros.com), 2026.
Run 7: more Greek speech
Trained on FLEURS Greek and English (as run 4) plus:
- Common Voice Greek, version 22 (Mozilla, CC0), through the
fsicoli/common_voice_22_0copy: the train, dev and unvalidated clips, with down-voted clips removed, at most 300 clips per speaker, and the test split's speakers and sentences removed; - STOMA Greek studio readings (A. Angelakis et al., CC BY 4.0), at most two readings of a sentence, with speaker M1 held out.
The audio was transcribed by RecMeets' Whisper and each word labelled by aligning it with the reference. No audio or text from these datasets is redistributed here, and no speaker is identified.
Held-out tests. AUROC, all words / unsure words, against run 4 on the same words (95% paired bootstrap interval):
| Test | Run 4 | Run 7 | Difference |
|---|---|---|---|
| Common Voice Greek (other speakers, their own microphones) | 0.773 / 0.863 | 0.843 / 0.908 | +0.070 [+0.048, +0.087] / +0.045 [+0.015, +0.078] |
| STOMA Greek (a held-out speaker) | 0.887 / 0.902 | 0.937 / 0.938 | +0.050 [+0.035, +0.064] / +0.036 [+0.001, +0.083] |
| FLEURS Greek | 0.904 / 0.905 | 0.899 / 0.889 | β0.005 [β0.013, +0.003] / β0.016 [β0.032, +0.004] |
| FLEURS English | 0.762 / 0.859 | 0.743 / 0.836 | β0.017 [β0.041, +0.004] / β0.023 [β0.057, +0.009] |
| AMI English meetings (noisy labels) | 0.578 / 0.608 | 0.541 / 0.619 |
Calibration (Brier) on FLEURS Greek is 0.090, against run 4's 0.085.
Caveats:
- It missed the publishing rule set before the night's runs: no drop over 0.01 on FLEURS unsure words.
- It was chosen from nine runs scored on these same tests, so its gains are slightly flattering.
- On real English meetings it stays close to chance, as run 4 does.
Attribution: FLEURS (Google, CC BY 4.0); Common Voice (Mozilla, CC0); STOMA (A. Angelakis et al., CC BY 4.0).
- Downloads last month
- -
Model tree for somniusx/recmeets-laya-words
Base model
convaiinnovations/laya