Scaling Automated Post-Training
- Type
- repo
- Venue
- IntologyAI (GitHub)
- Year
- 2026
- Source
- github
- Access
- free
- Added
- 2026-08-14T18:45:00Z
- Verified
- 2026-08-14T18:45:00Z
Summary
Figures-only companion to Intology blog Scaling Automated Post-Training. Locus (updated automated research system) scores 44.7 on official PostTrainBench vs Claude Code (Fable 5) 41.8; under PostTrainBench+ reaches 51.6% composite vs the human-tuned Qwen3-1.7B-Instruct checkpoint at 49.4%; AIME 2025 20% (double next baseline) with clearer scaling vs training-token count; more unique approaches per benchmark; higher peak average rank than other competitors on live prize-money Kaggle competitions with public leaderboards; Bubble production model ~2.8x lower error, ~5.4x lower latency, 105x lower cost vs legacy. Full solutions and artifacts marked to follow. MIT.
Keywords
locus · intology · posttrainbench · automated-post-training · kaggle · aime
Topics
automated post-training, LLM agents, AutoML
Research notes
- Primary: GitHub README + API (MIT, language null, 6 stars / 0 forks at check). Blog https://www.intology.ai/blog/scaling-automated-post-training. Discord posted the repo plus https://x.com/intology/status/2084319121332965804. Figures-only companion; solutions not released at check. Not a new hosted corpus, so no datasets_local row.