cdli/ugandan_luganda_nonstandard_speech_v1.0
Viewer โข Updated โข 8.14k โข 13
Fine-tuned versions of openai/whisper-large-v3 for automatic speech recognition of Luganda, trained on nonstandard speech from the CDLI Ugandan Luganda Nonstandard Speech dataset.
| Run | Folder | Decoder Frozen | Augmentation | SpecAugment | Best Checkpoint | Avg WER |
|---|---|---|---|---|---|---|
| Run 1 (v16) | v16-checkpoint-350 | Yes (encoder only) | Yes | Off | Step 350 | 53.4% |
| Run 2 (v17) | v17-checkpoint-350 | No (full model) | No | Off | Step 350 | 53.3% |
| Run 3 (v20) | v20-checkpoint-500 | No (full model) | No | On | Step 500 | 53.9% |
Preferred model: Run 2 (v17) โ best performance on moderate and severe speakers.
from transformers import pipeline
asr = pipeline(
"automatic-speech-recognition",
model="ElizabethMwangi/whisper-large-v3-luganda-nss",
model_kwargs={"subfolder": "v17-checkpoint-350"}
)
result = asr("audio.wav")
print(result["text"])
\```
## Model Details
- **Base model:** openai/whisper-large-v3
- **Language:** Luganda (`lg`)
- **Task:** Automatic Speech Recognition (ASR)
- **Training data:** cdli/ugandan_luganda_nonstandard_speech_v1.0
- **Language token:** sw (Swahili used as proxy for Luganda)
- **Length limit:** enabled
## Intended Use
This model is intended for transcription of Luganda nonstandard speech, including dysarthric, stuttering, and otherwise atypical speech patterns. It is part of ongoing research into inclusive ASR for low-resource Ugandan languages.
## Limitations
- Optimised for Luganda nonstandard speech; performance on standard Luganda or other languages may vary
- Language token is set to Swahili (`sw`) as a proxy for Luganda
## Citation
If you use this model, please cite the CDLI Ugandan Luganda Nonstandard Speech dataset and this repository.
Base model
openai/whisper-large-v3