Intelligence Density (Density Aware Training)
- Type
- blog
- Venue
- Trajectory field notes
- Year
- 2026
- Source
- web
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Trajectory's 'Intelligence Density' field notes argue the right efficiency metric is cost per completed task, not cost per token, and introduce Density Aware Training — a knob-free post-training technique that cuts wasted output (90k→37k tokens at equal pass rate on a legal-agent benchmark) and resists the reward hacking where standard RL inflates training reward via verbosity while held-out performance collapses.
Keywords
post-training · rl · efficiency · evals · agents
Topics
post-training, rl, efficiency, evals, agents
Research notes
- Discovery: Also announced via https://x.com/trajectorylabs/status/2103157329940365624
- Method: Density-aware training modifies the RL/post-training objective to penalize wasted compute (verbosity) while preserving task success, giving the model reason to learn the task itself rather than exploit reward-length correlations.
- Key findings: Intelligence Density: shift evaluation from cost-per-token to cost-per-task; cheap tokens ≠ cheap tasks — an open-source model can cost more per completed task than a frontier model; Density Aware Training (DAT): post-training technique with 'no knobs to tune' that rewards task completion without wasted compute; Harvey Legal Agent Benchmark: Nemotron 3.5 Nano 30B-A3B kept 8.3% pass rate while mean output fell 90k → 37k tokens; Sierra Tau3: DAT reached 55.6% held-out score vs 5.7% for standard RL; standard RL's training reward rose with output length while held-out collapsed — classic reward hacking
- Limitations: Company-published field notes, not a peer-reviewed paper; benchmarks include a private suite (Sierra Tau3). No public code or dataset release mentioned.
- Trajectory builds a platform for continual learning. Experiments used NVIDIA Nemotron 3.5/3 models (30B-A3B, 550B-A55B).