Running 2 When the Benchmark Is Just a Rubber Stamp ⚖ 2 How scaffolding alone scored 27% on legal review
Running 57 physics-intern: an Autonomous Agent for Physics Research 📝 57 Explore an autonomous AI workflow for physics research
Running Agents 12 Token Count Viewer ⚡ 12 Explore token counts across datasets with interactive charts
Running on CPU Upgrade 277 The Synthetic Data Playbook: Generating Trillions of the Finest Tokens 📝 277 Visualize synthetic‑data experiments as an interactive bookshelf
Running Featured 84 QED-Nano: Teaching a Tiny Model to Prove Hard Theorems 📝 84 Who needs 1T parameters? Olympiad proofs with a 4B model
Running 17 The Jagged AI Frontier is a Data Frontier 🧭 17 Why AI capabilities are shaped by data availability
Running Agents 107 Internal European Leaderboard 🌍 107 Explore and compare multilingual LLM benchmarks