← Back to explorer

Tmax: A simple recipe for terminal agents

Type
paper
Venue
arXiv
Year
2026
Source
huggingface
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Ai2's recipe for training terminal agents with reinforcement learning: the TMax-15K task corpus with verifiers and prebuilt environments, used to train the tmax model family on Qwen3 bases.

Research notes

  • Key findings: TMax-15K task corpus with pytest verifiers and prebuilt Docker environments as an RL training setup for terminal agents; Trained tmax model family (2b/4b/8b/9b/27b) on Qwen3 bases with DPPO
  • Paper record for the research behind the tmax-15k-open-instruct dataset.