DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Paper • 2504.11456 • Published • 12
A complete, reproducible pipeline to improve on GSM8K baselines using contamination-free publicly available data.
Key guarantees:
| Component | File | Purpose |
|---|---|---|
| Contamination Audit | contamination_audit.py |
Proves no test set leakage |
| Training Pipeline | train_gsm8k.py |
SFT training with Trackio monitoring |
| Ablation Study | ablation_study.py |
Sweeps LR, epochs, data subsets |
| Eval Harness | eval_harness.py |
Reproducible evaluation with CIs |
| Report Generator | reproducibility_report.py |
Signed final report |
meta-math/MetaMathQA) — 395K math problemsTwo independent methods (union of flagged samples excluded):
N-gram overlap (per Qwen2.5-Math, arXiv:2409.12122):
Embedding similarity (per DeepMath-103K, arXiv:2504.11456):
all-MiniLM-L6-v2| Variable | Values |
|---|---|
| Learning rate | 1e-5, 2e-5, 5e-5 |
| Epochs | 1, 2, 3 |
| Data subset | all, gsm_only, math_only |
Fixed: seed=42, LoRA r=32, batch_size=32 effective, max_seq_length=2048
| Model | GSM8K | Source |
|---|---|---|
| Qwen2.5-1.5B-Instruct (base) | ~50% | Model card |
| LLaMA-2-7B (no SFT) | 14.6% | arXiv:2309.12284 |
| MetaMath-7B | 66.5% | arXiv:2309.12284 |
| MetaMath-70B | 82.3% | arXiv:2309.12284 |
pip install transformers trl datasets peft torch scipy sentence-transformers trackio accelerate numpy
python contamination_audit.py
ABLATION_ID=lr2e5_ep2_all ABLATION_LR=2e-5 ABLATION_EPOCHS=2 ABLATION_SUBSET=all \
MAX_TRAIN_SAMPLES=50000 accelerate launch ablation_study.py
python train_gsm8k.py
MODEL_PATH=YOUSSEF88/gsm8k-ablation-lr2e5_ep2_all python eval_harness.py
python reproducibility_report.py
@misc{gsm8k_improvement_2026,
title={GSM8K Improvement with Contamination-Free Training},
author={YOUSSEF88},
year={2026},
note={Reproducible pipeline with contamination audit}
}
This model repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "YOUSSEF88/gsm8k-improvement-pipeline"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
For non-causal architectures, replace AutoModelForCausalLM with the appropriate AutoModel class.