← Back to explorer

GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Type
paper
Venue
arXiv / ICLR 2026
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T19:15:00Z
Verified
2026-08-14T19:15:00Z

Summary

GEPA (Genetic-Pareto; ICLR 2026 Oral). For a compound LLM system, sample traces, reflect with an LM on textual feedback (compiler errors, missed gold docs, constraint fails), mutate one module’s prompt, keep the candidate if the minibatch improves, and parent-sample from the instance-wise Pareto front instead of always mutating the global best. On Qwen3 8B across HotpotQA/IFBench/HoVer/PUPA/AIME-2025/LiveBench-Math: +9.62 pp vs baseline vs GRPO +3.68 pp at 24k rollouts, using ~4–35× fewer rollouts (aggregate ~3936). Also beats MIPROv2; GPT-4.1 Mini aggregate +12.19 pp. Code MIT gepa-ai/gepa; first-class dspy.GEPA. Discord post is Akshay Pachaar’s April 2026 X article “How to Beat GRPO Without Touching Model Weights” (DSPy cookbook; claims Berkeley beat GRPO by 10 points with 35× fewer rollouts).

Keywords

gepa · dspy · prompt-optimization · grpo · pareto · iclr-2026 · berkeley · x

Topics

prompt optimization, compound AI systems, DSPy

Research notes

  • Primary: arxiv abs 2507.19457. Discord/X article https://x.com/akshay_pachaar/status/2049916107923034300 via fxtwitter (empty tweet body; article id 2049896113206149120). Code https://github.com/gepa-ai/gepa (MIT). Docs https://dspy.ai/api/optimizers/GEPA/overview/. Artifact https://github.com/gepa-ai/gepa-artifact. Optimizer/code, not a hosted corpus, so no datasets_local row.