← Back to explorer

Reward design lessons from 7,780 MiMo RL environments (X thread)

Type
x_thread
Venue
X (Twitter)
Year
2026
Source
x
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

X thread of reward-design lessons learned from the 7,780 RL environments Xiaomi open-sourced for MiMo training. Thread starts from the premise that in RL environments, reward design is everything. Visible lessons include: verifying rewards, reward terms and cheating (gaming), scoring broken runs as 0, and a FineEnvs-built explorer for all 7,780 environments that lets anyone open a task to read the exact agent prompt, grader and judge prompts, then run a rollout with any Hugging Face model or a custom endpoint.

Keywords

RL · reward design · RL environments · MiMo · agents · reward hacking

Topics

reinforcement learning, reward design, RL environments

Research notes

  • Discovery: Adithya S K (@adithya_s_k, Scaling RL Envs @huggingface; prev MSFTResearch, Apple ML) on 2026-09-30: https://x.com/adithya_s_k/status/2105186691007017123 (quote-tweets his own 2026-09-26 post: https://x.com/adithya_s_k/status/2103737116857708849)
  • The environments: Xiaomi MiMo-V2.6-RL-oss, already cataloged in the Datasets spreadsheet. FineEnvs converted all 7,780 into Harbor tasks (code 2,698 / cyber 1,000 / general 925 / terminal 64 / webdev 2,093 / music 1,000), also cataloged in the Datasets spreadsheet.
  • Explorer: https://huggingface.co/spaces/FineEnvs/MiMo-RL-Envs-Explorer
  • Substantive research associated with the dataset; logged here per project rule.