GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Type
- paper
- Venue
- arXiv / ICLR 2026
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:15:00Z
- Verified
- 2026-08-14T19:15:00Z
Summary
GEPA (Genetic-Pareto; ICLR 2026 Oral). For a compound LLM system, sample traces, reflect with an LM on textual feedback (compiler errors, missed gold docs, constraint fails), mutate one module’s prompt, keep the candidate if the minibatch improves, and parent-sample from the instance-wise Pareto front instead of always mutating the global best. On Qwen3 8B across HotpotQA/IFBench/HoVer/PUPA/AIME-2025/LiveBench-Math: +9.62 pp vs baseline vs GRPO +3.68 pp at 24k rollouts, using ~4–35× fewer rollouts (aggregate ~3936). Also beats MIPROv2; GPT-4.1 Mini aggregate +12.19 pp. Code MIT gepa-ai/gepa; first-class dspy.GEPA. Discord post is Akshay Pachaar’s April 2026 X article “How to Beat GRPO Without Touching Model Weights” (DSPy cookbook; claims Berkeley beat GRPO by 10 points with 35× fewer rollouts).
Keywords
gepa · dspy · prompt-optimization · grpo · pareto · iclr-2026 · berkeley · x
Topics
prompt optimization, compound AI systems, DSPy
Research notes
- Primary: arxiv abs 2507.19457. Discord/X article https://x.com/akshay_pachaar/status/2049916107923034300 via fxtwitter (empty tweet body; article id 2049896113206149120). Code https://github.com/gepa-ai/gepa (MIT). Docs https://dspy.ai/api/optimizers/GEPA/overview/. Artifact https://github.com/gepa-ai/gepa-artifact. Optimizer/code, not a hosted corpus, so no datasets_local row.