TMax

💻 Code · 🤗 Models & Data · 📜 Paper · 📓 Blog

TMax v1.1 9B

This terminal-agent model was trained using DPPO on Qwen/Qwen3.5-9B with the TMax hardened prefilter task set.

This model is part of the TMax v1.1 collection. The original TMax recipe is described in our paper.

Checkpoints

This branch contains step 200. The main branch contains step 200, the highest-scoring available checkpoint in the completed TB-Lite sweep. Checkpoint selection uses TB-Lite, not Terminal-Bench 2.1.

Checkpoints 100, 200, 300, 400, 500 are provided as branches named step-0100, step-0200, etc.

Evaluation Results

Training step TB-Lite (%) ± SE (pp) Selection
100 65.24 ± 1.81
200 65.62 ± 2.03 main
300 62.81 ± 1.89
400 63.30 ± 2.11
500 59.02 ± 1.67

The selected checkpoint scores 29.55 ± 1.78% on Terminal-Bench 2.1 (88 tasks, three trials per task; configure-git-webserver excluded).

These v1.1 scores use the fixed Vanillux0.2.3 harness with Sandfleet through Harbor, three trials per task (96 TB-Lite tasks). Scores and standard errors are taken from the completed evaluation sweep updated on October 5, 2026. The harness differs from the original release; use the original model cards for v1.0 results. The per-checkpoint results and pinned source revisions are included in release_manifest.json.

Model Details

Model Description

Use

Serve the model with vLLM and use the terminal-agent harness in our codebase:

uvx vllm==0.19.1 serve hamishivi/tmax-v1.1-9b \
  --served-model-name tmax-v1.1 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --port 8008 \
  --max-model-len 65536 \
  --tensor-parallel-size 8

The Qwen3.5/3.6 checkpoints contain the text-only causal language model; use the Qwen3 XML tool parser.

Training Details

  • Released training step: 200
  • Main checkpoint: 200, selected on TB-Lite
  • Available milestone checkpoints: 100, 200, 300, 400, 500

The source checkpoint is TMaxxx/qwen35-9b-dppo-hardened-prefilter at step 200. Original checkpoint files are preserved, including campaign_provenance.json where supplied. See release_manifest.json for the immutable source commits and evaluation results. See the training launch scripts and original provenance for run-specific settings.

License

This model is licensed under Apache 2.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.

Citation

If you use our model or data, please cite our paper:

@misc{ivison2026tmaxsimplerecipeterminal,
      title={Tmax: A simple recipe for terminal agents}, 
      author={Hamish Ivison and Junjie Oscar Yin and Rulin Shao and Teng Xiao and Nathan Lambert and Hannaneh Hajishirzi},
      year={2026},
      eprint={2606.23321},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2606.23321}, 
}
Downloads last month
30
Safetensors
Model size
9B params
Tensor type
BF16
·
Video Preview
loading

Model tree for hamishivi/tmax-v1.1-9b

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(997)
this model
Quantizations
2 models

Dataset used to train hamishivi/tmax-v1.1-9b

Paper for hamishivi/tmax-v1.1-9b