← Back to explorer

Intelligence Density (Density Aware Training)

Type
blog
Venue
Trajectory field notes
Year
2026
Source
web
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Trajectory's 'Intelligence Density' field notes argue the right efficiency metric is cost per completed task, not cost per token, and introduce Density Aware Training — a knob-free post-training technique that cuts wasted output (90k→37k tokens at equal pass rate on a legal-agent benchmark) and resists the reward hacking where standard RL inflates training reward via verbosity while held-out performance collapses.

Keywords

post-training · rl · efficiency · evals · agents

Topics

post-training, rl, efficiency, evals, agents

Research notes

  • Discovery: Also announced via https://x.com/trajectorylabs/status/2103157329940365624
  • Method: Density-aware training modifies the RL/post-training objective to penalize wasted compute (verbosity) while preserving task success, giving the model reason to learn the task itself rather than exploit reward-length correlations.
  • Key findings: Intelligence Density: shift evaluation from cost-per-token to cost-per-task; cheap tokens ≠ cheap tasks — an open-source model can cost more per completed task than a frontier model; Density Aware Training (DAT): post-training technique with 'no knobs to tune' that rewards task completion without wasted compute; Harvey Legal Agent Benchmark: Nemotron 3.5 Nano 30B-A3B kept 8.3% pass rate while mean output fell 90k → 37k tokens; Sierra Tau3: DAT reached 55.6% held-out score vs 5.7% for standard RL; standard RL's training reward rose with output length while held-out collapsed — classic reward hacking
  • Limitations: Company-published field notes, not a peer-reviewed paper; benchmarks include a private suite (Sierra Tau3). No public code or dataset release mentioned.
  • Trajectory builds a platform for continual learning. Experiments used NVIDIA Nemotron 3.5/3 models (30B-A3B, 550B-A55B).