GSM8K Improvement Pipeline

Overview

A complete, reproducible pipeline to improve on GSM8K baselines using contamination-free publicly available data.

Key guarantees:

  • ✅ No benchmark contamination (dual-method audit: 13-gram + embedding similarity)
  • ✅ No GSM8K test leakage in training data
  • ✅ No synthetic CoT derived from GSM8K test set
  • ✅ Bootstrap 95% confidence intervals on all results
  • ✅ Signed reproducibility report with exact seeds and hashes

Architecture

Component File Purpose
Contamination Audit contamination_audit.py Proves no test set leakage
Training Pipeline train_gsm8k.py SFT training with Trackio monitoring
Ablation Study ablation_study.py Sweeps LR, epochs, data subsets
Eval Harness eval_harness.py Reproducible evaluation with CIs
Report Generator reproducibility_report.py Signed final report

Method

Base Model

  • Qwen/Qwen2.5-1.5B-Instruct — strong math pretraining, efficient size

Training Data

  • MetaMathQA (meta-math/MetaMathQA) — 395K math problems
    • Derived from GSM8K and MATH training sets only (never test)
    • Augmentation types: AnsAug, Rephrased, SV, FOBAR
    • Paper: MetaMath (arXiv:2309.12284)

Training Method

  • SFT with LoRA (r=32, alpha=64, all-linear)
  • Sequence packing for efficiency
  • Cosine LR schedule with warmup

Contamination Audit Protocol

Two independent methods (union of flagged samples excluded):

  1. N-gram overlap (per Qwen2.5-Math, arXiv:2409.12122):

    • 13-gram matching between train questions and GSM8K test questions
    • LCS ratio > 0.6 threshold for confirmation
  2. Embedding similarity (per DeepMath-103K, arXiv:2504.11456):

    • Sentence embeddings via all-MiniLM-L6-v2
    • Cosine similarity > 0.95 threshold

Ablation Design

Variable Values
Learning rate 1e-5, 2e-5, 5e-5
Epochs 1, 2, 3
Data subset all, gsm_only, math_only

Fixed: seed=42, LoRA r=32, batch_size=32 effective, max_seq_length=2048

Baselines (from literature)

Model GSM8K Source
Qwen2.5-1.5B-Instruct (base) ~50% Model card
LLaMA-2-7B (no SFT) 14.6% arXiv:2309.12284
MetaMath-7B 66.5% arXiv:2309.12284
MetaMath-70B 82.3% arXiv:2309.12284

Running the Pipeline

Prerequisites

pip install transformers trl datasets peft torch scipy sentence-transformers trackio accelerate numpy

Step 1: Contamination Audit

python contamination_audit.py

Step 2: Run Ablation (single config)

ABLATION_ID=lr2e5_ep2_all ABLATION_LR=2e-5 ABLATION_EPOCHS=2 ABLATION_SUBSET=all \
  MAX_TRAIN_SAMPLES=50000 accelerate launch ablation_study.py

Step 3: Run Full Training (best config)

python train_gsm8k.py

Step 4: Standalone Evaluation

MODEL_PATH=YOUSSEF88/gsm8k-ablation-lr2e5_ep2_all python eval_harness.py

Step 5: Generate Report

python reproducibility_report.py

Hardware Requirements

  • Training: NVIDIA A10G (24GB) or better
  • Evaluation: Same (for fast inference)
  • Contamination audit: CPU only (8GB+ RAM)
  • Estimated runtime: ~3-4h per ablation config

Citation

@misc{gsm8k_improvement_2026,
  title={GSM8K Improvement with Contamination-Free Training},
  author={YOUSSEF88},
  year={2026},
  note={Reproducible pipeline with contamination audit}
}

References

Generated by ML Intern

This model repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "YOUSSEF88/gsm8k-improvement-pipeline"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)

For non-causal architectures, replace AutoModelForCausalLM with the appropriate AutoModel class.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Papers for YOUSSEF88/gsm8k-improvement-pipeline