← Back to explorer

Extending SWE-2's Reward Function to Steer the Pareto Frontier

Type
paper
Venue
Blog write-up, anishlk.com, 25 Sep 2026
Year
2026
Source
web
Access
free
Language
English
Added
2026-10-02
Verified
2026-10-02

Summary

Extension of Cognition's SWE-2 RL reward function R = S - lambda^(e) C. SWE-2's 'slope-matched penalty' sets lambda^(e) to the slope of the base model's Pareto curve, which pushes the frontier but is unopinionated about where on the frontier the model lands (all points on an iso-reward line earn equal reward). This work steers the direction of improvement with a tradeoff parameter alpha in [0,1]: maximizing tau subject to u_e(pi) >= alpha*tau and v_e(pi) >= (1-alpha)*tau, where u_e, v_e are relative success/cost improvements. Reformulating max_pi min_beta of a weighted Lagrangian gives the same SWE-2-form reward with lambda as a function of beta, yielding the per-RL-step update log lambda^(e) <- log lambda_prev^(e) + eta[(1-alpha) u_e_hat - alpha v_e_hat]: raise lambda when success gains exceed the alpha mix, lower it when cost savings do. On a constructed toy family of effort-adjustable RL tasks (100 problems, 5 trainable parameters, closed-form Pareto curve), each alpha's median improvement direction lands within 1 degree of its target; fixed lambda increases both success and cost.

Keywords

SWE-2 · Pareto frontier · reward function · adaptive lambda · RL post-training · cost-performance tradeoff

Topics

SWE-2, Pareto frontier, reward shaping, RL post-training, cost-performance tradeoff, adaptive lambda

Research notes

  • Discovery: @_anishlk (Anish Lakkapragada) X thread 2026-10-02 1:41 PM ET (https://x.com/_anishlk/status/2106077176630313318); Silas Alberti (Cognition) replied 'Nice work!'
  • Full derivations on anishlk.com/swe-2-extended/ (dated Sep 25, 2026); explainer video embedded; code repo linked from the write-up
  • Experiments use the write-up's code, 1,000 training steps per problem, 3 seeds
  • Video style inspired by @trajectorylabs' recent posts, visuals/sound by Opus 5.5
  • Early readers: @marsxiang, @neilkale, @_rexliu, @rrebeccajjoseph, @marcmelikyan, @kento_nishi
  • Builds on Cognition's SWE-2 blog (cognition.com) and FrontierCode 1.1