← Back to explorer

Synthetic Computers at Scale for Long-Horizon Productivity Simulation

Type
other
Venue
arXiv / Microsoft

Summary

Builds user-specific computers (persona → profile → filesystem plan → content-rich artifacts) then runs month-scale sims: a setup agent writes deliverables, a work agent (Claude Sonnet 4.6; Opus 4.6 for setup) executes over the machine with simulated collaborators. 1,000 computers; mean 2,272 turns / 8.59h per run. Occupation skills from 900 train sims lift held-out rubric 61.6%→68.6% (win 83/100) and transfer to GDPVal (Sonnet 105 wins / 67 losses). Public release: 100 computers plus 500 retrospective reports (HF 98 rows).

Keywords

synthetic-computers · agents · long-horizon · productivity · gdpval · persona · microsoft · claude

Topics

LLM agents, synthetic environments, long-horizon productivity

Research notes

  • Primary: arxiv abs (cs.AI; also cs.CL, cs.LG). CC BY 4.0 on HTML. Microsoft. Correspondence {taoge, baolinpeng}@microsoft.com. HF dataset microsoft/synthetic-computers-at-scale (MIT, 98 computers, 19 likes / 176 downloads at check). HF paper page 21 upvotes; 1 linked dataset not copied into hf_* fields. Discord posted abs. Substantial public release → datasets_local row. License field left blank per catalog convention.