← Back to explorer

ReZero

Type
repo
Venue
Menlo Research (GitHub)
Year
2026
Source
github
Access
free
Added
2026-08-14T18:35:00Z
Verified
2026-08-14T18:35:00Z

Summary

ReZero (Retry-Zero) adds a reward_retry term to GRPO so the policy is paid for issuing a later search query after an unsuccessful first attempt, but only if the trajectory still emits a well-formed final answer. Paper arXiv 2504.11001: Llama-3.2-3B-Instruct on Apollo 3 mission chunks (341 chunks / 32 held-out) peaks at 46.88% accuracy vs 25% without the retry reward (1xH200). Repo trains via train_grpo.py, ships data/ plus a Gradio demo, and points to Menlo/ReZero-v0.1-llama-3.2-3b-it-grpo-250404 (and GGUF). GitHub menloresearch/ReZero redirects to janhq/ReZero. No LICENSE file on the repo; arXiv Atom API states no license.

Keywords

rezero · rag · grpo · retry · search · menlo · llama-3.2 · retrieval

Topics

RAG, reinforcement learning, search agents

Research notes

  • Primary: GitHub README + API (no SPDX license; 160 stars / 9 forks at check; owner now janhq, posted URL menloresearch/ReZero). Paper arXiv 2504.11001 (cs.CL; Atom API has no license). HF weights Menlo/ReZero-v0.1-llama-3.2-3b-it-grpo-250404 not copied into hf_* because the leftover is the GitHub repo. Discord posted the repo. Apollo training chunks in data/ are a small in-repo split, not a new hosted corpus, so no datasets_local row.