Reward design lessons from 7,780 MiMo RL environments (X thread)
- Type
- x_thread
- Venue
- X (Twitter)
- Year
- 2026
- Source
- x
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
X thread of reward-design lessons learned from the 7,780 RL environments Xiaomi open-sourced for MiMo training. Thread starts from the premise that in RL environments, reward design is everything. Visible lessons include: verifying rewards, reward terms and cheating (gaming), scoring broken runs as 0, and a FineEnvs-built explorer for all 7,780 environments that lets anyone open a task to read the exact agent prompt, grader and judge prompts, then run a rollout with any Hugging Face model or a custom endpoint.
Keywords
RL · reward design · RL environments · MiMo · agents · reward hacking
Topics
reinforcement learning, reward design, RL environments
Research notes
- Discovery: Adithya S K (@adithya_s_k, Scaling RL Envs @huggingface; prev MSFTResearch, Apple ML) on 2026-09-30: https://x.com/adithya_s_k/status/2105186691007017123 (quote-tweets his own 2026-09-26 post: https://x.com/adithya_s_k/status/2103737116857708849)
- The environments: Xiaomi MiMo-V2.6-RL-oss, already cataloged in the Datasets spreadsheet. FineEnvs converted all 7,780 into Harbor tasks (code 2,698 / cyber 1,000 / general 925 / terminal 64 / webdev 2,093 / music 1,000), also cataloged in the Datasets spreadsheet.
- Explorer: https://huggingface.co/spaces/FineEnvs/MiMo-RL-Envs-Explorer
- Substantive research associated with the dataset; logged here per project rule.