InstructLR: A Scalable Approach to Create Instruction Dataset for Under-Resourced Languages Paper • 2512.02213 • Published Dec 1, 2025 • 3
view article Article YODAS v3: A 1 Million Hour Dataset for the Next Generation of Open Voice AI Research espnet • 14 days ago • 36
view article Article How to Fine-Tune Nemotron 3.5 ASR for Your Language, Domain, or Accent nvidia • Jun 4 • 80
BidirLM Collection BidirLM is a family of 5 frontier bidirectional encoders, including an omnimodal variant at 2.5B. • 8 items • Updated Apr 15 • 4
view article Article DenseOn with the LateOn: Open State-of-the-Art Single and Multi-Vector Models lightonai • Apr 21 • 46
view article Article Fine-Tune W2V2-Bert for low-resource ASR with 🤗 Transformers ylacombe • Jan 19, 2024 • 50
Embarrassingly Simple Self-Distillation Improves Code Generation Paper • 2604.01193 • Published Apr 1 • 56
TRACER: Trace-Based Adaptive Cost-Efficient Routing for LLM Classification Paper • 2604.14531 • Published Apr 16 • 9
view article Article How I contributed a new model to the Transformers library using Codex nielsr • Mar 30 • 53
fiNERweb Collection A multilingual dataset for NER covering 91 langauges and 25 scripts • 3 items • Updated Dec 16, 2025 • 3
Fine-tune ready versions of the LLMSQL benchmark Collection Fine-tune-ready (0/1/5-shot) version of LLMSQL 1.0. LLMSQL 2.0 is a test-only benchmark and has no training data. • 1 item • Updated 1 day ago • 1
Beyond Language Modeling: An Exploration of Multimodal Pretraining Paper • 2603.03276 • Published Mar 3 • 109