Tmax: A simple recipe for terminal agents
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- huggingface
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Ai2's recipe for training terminal agents with reinforcement learning: the TMax-15K task corpus with verifiers and prebuilt environments, used to train the tmax model family on Qwen3 bases.
Research notes
- Key findings: TMax-15K task corpus with pytest verifiers and prebuilt Docker environments as an RL training setup for terminal agents; Trained tmax model family (2b/4b/8b/9b/27b) on Qwen3 bases with DPPO
- Paper record for the research behind the tmax-15k-open-instruct dataset.