← Back to explorer

TMax-15k Open Instruct

Type
dataset
Venue
Allen Institute for AI (Ai2)
Year
2026
Source
huggingface
Access
free
Language
English
Added
2026-07-17T20:18:03.829244+00:00
Verified
2026-07-17T20:18:03.829244+00:00

Summary

TMax-15k is a dataset of ~14,600 reinforcement learning environment instances for training terminal-based coding agents, generated through a synthetic data pipeline using a frontier model with novel taxonomy, difficulty control, personas, and verifier diversification. It is over 2.5x larger than previously released terminal-agent datasets and significantly harder. The dataset was used to train TMax 9B, which achieves 27% on Terminal-Bench 2.0 using a simple outcome-only RL recipe, outperforming much larger models.

Keywords

code terminal-agents reinforcement-learning software-engineering sft rlvr tmax ai2 synthetic

Topics

Code / Software Engineering

Research notes

  • Associated with paper 'TMax: A simple recipe for terminal agents' (arxiv 2606.23321). Formatted for use with Ai2's open-instruct framework. Tasks involve bash terminal exploration, codebase understanding, and solution implementation.