ReZero
- Type
- repo
- Venue
- Menlo Research (GitHub)
- Year
- 2026
- Source
- github
- Access
- free
- Added
- 2026-08-14T18:35:00Z
- Verified
- 2026-08-14T18:35:00Z
Summary
ReZero (Retry-Zero) adds a reward_retry term to GRPO so the policy is paid for issuing a later search query after an unsuccessful first attempt, but only if the trajectory still emits a well-formed final answer. Paper arXiv 2504.11001: Llama-3.2-3B-Instruct on Apollo 3 mission chunks (341 chunks / 32 held-out) peaks at 46.88% accuracy vs 25% without the retry reward (1xH200). Repo trains via train_grpo.py, ships data/ plus a Gradio demo, and points to Menlo/ReZero-v0.1-llama-3.2-3b-it-grpo-250404 (and GGUF). GitHub menloresearch/ReZero redirects to janhq/ReZero. No LICENSE file on the repo; arXiv Atom API states no license.
Keywords
rezero · rag · grpo · retry · search · menlo · llama-3.2 · retrieval
Topics
RAG, reinforcement learning, search agents
Research notes
- Primary: GitHub README + API (no SPDX license; 160 stars / 9 forks at check; owner now janhq, posted URL menloresearch/ReZero). Paper arXiv 2504.11001 (cs.CL; Atom API has no license). HF weights Menlo/ReZero-v0.1-llama-3.2-3b-it-grpo-250404 not copied into hf_* because the leftover is the GitHub repo. Discord posted the repo. Apollo training chunks in data/ are a small in-repo split, not a new hosted corpus, so no datasets_local row.