← Back to explorer

PostTrainBench: Can LLM Agents Automate LLM Post-Training?

Type
paper
Venue
arXiv (cs.SE)
Year
2026
Source
huggingface
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Introduces PostTrainBench, a benchmark where CLI agents receive a base LLM, an evaluation script, and 10 H100-hours to autonomously improve the model via any post-training strategy, testing whether LLM agents can automate LLM post-training.

Research notes

  • Key findings: Proposes the 'agent post-trains the model' evaluation setup: base LLM + eval script + 10 H100-hours, any strategy allowed
  • Paper record for the research behind the PostTrainBench-Trajectories dataset.