Qwen3-0.6B 路 STaR round-1 (math reasoning)
Correctness-gated self-training (STaR / rejection-sampling SFT), round 1, applied to a Qwen3-0.6B math reasoner. Part of the qwen3-nano-math-reasoner post-training comparison.
- Seed: Qwen3-0.6B SFT-distilled from Qwen3-235B-A22B reasoning traces
(
jeetganatra/qwen3-0.6b-qwen3-235b-a22b-math-distill). - Method: sample k=4 rollouts/problem over 800 MATH problems (the held-out 500 are
excluded), keep only
\boxed{}-verified-correct rollouts (1,226 kept from 660/800 problems solved), then SFT on that filtered set.
Results (MATH-500, greedy)
| Model | Eval | Accuracy |
|---|---|---|
| SFT seed | full-500 @ 4096 | 37.8% (189/500) |
| STaR round-1 (this model) | full-500 @ 4096 | 36.0% (180/500) |
Round-1 STaR holds at the seed level (within the ~3 pp two-sample noise floor at
n=500), unlike on-policy distillation (OPD/OPSD) which degraded the same seed sharply
(to 18% / 16% on a 50-problem subset where the seed scores 50%). At this 0.6B capacity
ceiling, correctness-gating prevents regression but does not lift accuracy. Full
write-up: ON_POLICY_EXPERIMENTS.md in the GitHub repo.
Weights
Raw PyTorch state dict qwen3-0.6B-star-round1.pth (same format as the seed). Evaluate
with the repo's grader:
python src/evaluate_math500.py --runtime hf --which_model reasoning \
--checkpoint_path qwen3-0.6B-star-round1.pth \
--dataset_size 500 --max_new_tokens 4096 --eval_batch_size 16
Model tree for jeetganatra/qwen3-0.6b-math-star-round1
Base model
Qwen/Qwen3-0.6B-Base