Reducing Doom Loops with Final Token Preference Optimization
- Type
- other
- Venue
- Liquid AI Blog
- Year
- 2026
- Source
- web
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:06:00Z
- Verified
- 2026-08-14T19:06:00Z
Summary
Antidoom mines the first token of a detected loop (≥4 repeats, ≥60 chars), then trains LoRA with Final Token Preference Optimization (FTPO): one rejected token vs up to 20 chosen alternatives, logit-space KL, two-part regularization. Early LFM2.5-2.6B: 10.2%→1.4% doom-loop rate; Qwen3.5-4B greedy: 22.9%→1%, with eval gains attributed to fewer loops. Early-stop at chosen_win≈0.35. Prompt mix LiquidAI/antidoom-mix-v1.0 (478,229 rows) used to elicit loops.
Keywords
antidoom · ftpo · doom-loop · liquid-ai · lfm · qwen · repetition · blog · x
Topics
reasoning models, repetition, preference optimization
Research notes
- Primary: Liquid AI blog (cite liquidAI2026Antidoom). Discord posted https://t.co/zPFEt8lMdr plus https://x.com/liquidai/status/2074494130126811473. Code https://github.com/Liquid4All/antidoom (Apache-2.0). Substantial prompt mix → datasets_local Antidoom Mix v1.0. Written by Sam Paech with listed contributors.