The ultimate guide to multi-harness RL
- Type
- article
- Venue
- Hugging Face Space, 2026-10-01
- Year
- 2026
- Source
- article
- Access
- public
- Language
- en
- Added
- 2026-10-01
- Verified
- 2026-10-01
Summary
The ultimate guide to multi-harness RL (Adithya S Kolavi / FineEnvs, Hugging Face Space, 2026-10-01). Not a paper: a 33-figure ~60-page interactive research article plus full artifact release documenting how to do RL across multiple agent harnesses without modifying them -- claiming the first public multi-harness agent RL pipeline. Core idea: a 'capture proxy' sits between the agent harness and vLLM, speaks the four API formats coding agents use (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, Gemini), records exact token ids/logprobs sampled by vLLM, and reconstructs TRL-trainable traces, so the same rollouts run unmodified through 10+ harnesses. Full GRPO runs: LFM2.5-2.6B trained with multi-harness RL on 1,000 SmolDataEnvs data-analysis tasks, evaluated on 250 held-out tasks under four harnesses: base 42.2% -> multi-harness RL 54.2% pass@1 overall (step 1,000; best checkpoint step 700: 54.6%); per-harness OpenCode 49.6%, Claude Code 48.8%, Codex 53.6%, Mini-SWE-Agent 64.8%. OpenCode-only RL (step 900) reached 52.4% overall and won on its home harness (56.0%) but lost to multi-harness RL on Claude Code and Codex. Tool efficiency: correctness-only reward keeps tool calls high; adding a small bonus for fewer calls cut tool calls 31.1% on 356 task/harness pairs solved by both base and model, savings in all four harnesses. SFT comparison: fine-tuning on 3,189 successful Qwen3.8-27B teacher rollouts (17,929 examples, 2 epochs) plateaued -- OpenCode SFT 47.5% (-8.5% calls), multi-harness SFT 43.1% (-24.2% calls), both below RL (54.6%), showing multi-harness RL beats imitation at both performance and efficiency. Setup: async GRPO (TRL Async GRPO), 1,000 steps, reward = correctness x (1 + 0.1x15/(15+tool_calls)), 8 rollouts per GRPO group (one harness each, rotating), on 2xH100 (~32h OpenCode-only, 46h multi-harness); rollouts in E2B sandboxes via OpenEnv + Harbor. OpenEnv is BSD-3-Clause (PR #1036 capture proxy + Harbor, PR #1280 typed TrainingTrace API); TRL PR #6947 HarnessRolloutWorker (on TRL main, not v1.14.1). Trained models (full weights) under LFM Open License v1.0: LFM2.5-2.6B-multiharness-RL, -opencode-RL, -multiharness-SFT, -opencode-SFT, Qwen3.5-2B variants. Caveat from the model card: best-checkpoint selection used the test set (no separate validation set); single-run results on fixed harness versions.
Keywords
multi-harness RL · agent harnesses · GRPO · TRL · OpenEnv · vLLM · capture proxy · Claude Code · Codex · OpenCode · tool efficiency · SFT vs RL · SmolDataEnvs · Harbor · agentic RL
Topics
multi-harness RL, agent harnesses, GRPO, TRL, OpenEnv, vLLM, capture proxy, Claude Code, Codex, OpenCode, tool efficiency, SFT vs RL, SmolDataEnvs, Harbor
Research notes
- Discovery: X post by Adithya S Kolavi (@adithya_s_k, 2026-10-01 11:43 EDT): 'Excited to release The ultimate guide to multi-harness RL...'
- Collection: huggingface.co/collections/FineEnvs/smoldataenvs-multi-harness-rl-6abdfaaa8d74dacd481d5212 (21 items)
- Models: huggingface.co/FineEnvs/LFM2.5-2.6B-multiharness-RL (main), -opencode-RL, -multiharness-SFT, -opencode-SFT, Qwen3.5-2B variants
- Code: github.com/huggingface/OpenEnv (BSD-3-Clause), github.com/adithya-s-k/FineEnvs
- Demo: huggingface.co/spaces/FineEnvs/smoldataenv-multi-harness-harbor
- BibTeX @misc{kolavi2026multiharnessrl} on the page
- Prior FineEnvs work cataloged: MiMo-V2.6-RL-harbor splits (Datasets rows 74-79)
- Highly relevant to user's post-training/RL work.
- Per-harness results (via @ben_burtenshaw post): base LFM2.5-2.6B solves 62% of held-out tasks in mini-swe-agent vs 33% in claude code; after multi-harness RL the 4-harness average goes 42%->54% and claude code 33%->49%. Training enabled by OpenEnv capture proxy (OpenEnv PR #1036) recording exact token ids/logprobs per call.
- 2026-10-02 @huggingface repost (https://x.com/huggingface/status/2106034221005312448, 2:51 PM ET, ~17.4K views): framing — same model, same weights scores 62% in one agent harness and 33% in another. Proxy detail: don't touch the harness; point it at a proxy that speaks all four API formats (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, Gemini), records exact token ids and logprobs vLLM sampled, train on that — zero changes to Claude Code, Codex, OpenCode. Results: LFM2.5-2.6B trained across 4 harnesses 42%->54% with 31% fewer tool calls (bonus for solving in fewer steps); single-harness OpenCode training 34%->58% but multi-harness model improved everywhere. Imitation shortcut (3,189 Qwen3.8-27B rollouts) plateaued at 47.5%, below both RL runs. Infographic stats: 12 harnesses, 222 calls intercepted, 454,530 tokens captured. Everything open: OpenEnv capture proxy, TRL trainer, tasks, SFT data, training code, all seven trained models.