HuggingFaceFW/fineweb-edu
Viewer • Updated • 3.5B • 397k • 1.32k
Mostly Olmo-3 architecture 10M model with 2:1 SWA:GQA and HoPE embeddings. Trained on ~15B tokens of Fineweb-Edu (filtered to the splits before ChatGPT's release), DCLM-Baseline, and UltraData-Math.
| Tasks | Version | Filter | n-shot | Metric | Value | Stderr | ||
|---|---|---|---|---|---|---|---|---|
| arc_challenge | 1 | none | 0 | acc | ↑ | 0.1792 | ± | 0.0112 |
| none | 0 | acc_norm | ↑ | 0.2125 | ± | 0.0120 | ||
| arc_easy | 1 | none | 0 | acc | ↑ | 0.3695 | ± | 0.0099 |
| none | 0 | acc_norm | ↑ | 0.3552 | ± | 0.0098 | ||
| hellaswag | 1 | none | 0 | acc | ↑ | 0.2663 | ± | 0.0044 |
| none | 0 | acc_norm | ↑ | 0.2706 | ± | 0.0044 | ||
| piqa | 1 | none | 0 | acc | ↑ | 0.5604 | ± | 0.0116 |
| none | 0 | acc_norm | ↑ | 0.5675 | ± | 0.0116 |
Category N Raw Normalized
----------------------------------------------------------------------------------------------
elementary_school_math_continuation::addition::grades_1_2::easy 128 20.31% 20.31%
elementary_school_math_continuation::comparison::grades_2_3::medium 44 18.18% 18.18%
elementary_school_math_continuation::comparison_difference::grades_2_3::medium 48 39.58% 39.58%
elementary_school_math_continuation::data::grades_2_3::easy 43 20.93% 20.93%
elementary_school_math_continuation::division::grades_3_4::medium 54 35.19% 33.33%
elementary_school_math_continuation::fractions_counting::grades_3_4::medium 50 22.00% 24.00%
elementary_school_math_continuation::geometry_area::grades_4_5::medium 52 61.54% 61.54%
elementary_school_math_continuation::geometry_perimeter::grades_4_5::medium 45 35.56% 35.56%
elementary_school_math_continuation::measurement::grades_2_3::easy 76 31.58% 31.58%
elementary_school_math_continuation::money::grades_3_4::medium 64 28.12% 28.12%
elementary_school_math_continuation::multiplication::grades_3_4::medium 74 39.19% 40.54%
elementary_school_math_continuation::patterns::grades_3_4::medium 53 30.19% 30.19%
elementary_school_math_continuation::subtraction::grades_1_2::easy 117 28.21% 28.21%
elementary_school_math_continuation::time::grades_2_3::easy 55 94.55% 94.55%
elementary_school_math_continuation::two_step_add_subtract::grades_2_3::medium 46 19.57% 19.57%
elementary_school_math_continuation::two_step_addition::grades_2_3::medium 19 15.79% 15.79%
elementary_school_math_continuation::two_step_subtraction::grades_2_3::medium 32 28.12% 31.25%
====================================================================
../step25178_hf/ (9,835,712 params) RESULTS
====================================================================
Raw continuation accuracy 33.30%
Length-normalized accuracy 33.50%
Primary (acc_norm) 33.50%
====================================================================
(^ realistically, within stderr)
BananaMind Base Bench 1.1
Overall Elo: 890
Accuracy: 132/350 (37.71%)
Weighted accuracy: 34.26%
language_completion: Elo 1134 | 39/50 (78.00%) | weighted 78.98%
commonsense: Elo 789 | 17/50 (34.00%) | weighted 28.39%
world_knowledge: Elo 899 | 21/50 (42.00%) | weighted 42.32%
context_tracking: Elo 801 | 14/50 (28.00%) | weighted 24.84%
quantitative: Elo 861 | 14/50 (28.00%) | weighted 25.87%
logical_reasoning: Elo 983 | 19/50 (38.00%) | weighted 34.48%
code_completion: Elo 805 | 8/50 (16.00%) | weighted 16.24%
(^ yeouch)
@misc{olmo2026olmo3,
title={Olmo 3},
author={Team Olmo and : and Allyson Ettinger and Amanda Bertsch and Bailey Kuehl and David Graham and David Heineman and Dirk Groeneveld and Faeze Brahman and Finbarr Timbers and Hamish Ivison and Jacob Morrison and Jake Poznanski and Kyle Lo and Luca Soldaini and Matt Jordan and Mayee Chen and Michael Noukhovitch and Nathan Lambert and Pete Walsh and Pradeep Dasigi and Robert Berry and Saumya Malik and Saurabh Shah and Scott Geng and Shane Arora and Shashank Gupta and Taira Anderson and Teng Xiao and Tyler Murray and Tyler Romero and Victoria Graf and Akari Asai and Akshita Bhagia and Alexander Wettig and Alisa Liu and Aman Rangapur and Chloe Anastasiades and Costa Huang and Dustin Schwenk and Harsh Trivedi and Ian Magnusson and Jaron Lochner and Jiacheng Liu and Lester James V. Miranda and Maarten Sap and Malia Morgan and Michael Schmitz and Michal Guerquin and Michael Wilson and Regan Huff and Ronan Le Bras and Rui Xin and Rulin Shao and Sam Skjonsberg and Shannon Zejiang Shen and Shuyue Stella Li and Tucker Wilde and Valentina Pyatkin and Will Merrill and Yapei Chang and Yuling Gu and Zhiyuan Zeng and Ashish Sabharwal and Luke Zettlemoyer and Pang Wei Koh and Ali Farhadi and Noah A. Smith and Hannaneh Hajishirzi},
year={2026},
eprint={2512.13961},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2512.13961},
}
@misc{chen2024hopenovelpositionalencoding,
title={HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation},
author={Yuhan Chen and Ang Lv and Jian Luan and Bin Wang and Wei Liu},
year={2024},
eprint={2410.21216},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2410.21216},
}