← Back to explorer

Reducing Doom Loops with Final Token Preference Optimization

Type
other
Venue
Liquid AI Blog
Year
2026
Source
web
Access
free
Language
en
Added
2026-08-14T19:06:00Z
Verified
2026-08-14T19:06:00Z

Summary

Antidoom mines the first token of a detected loop (≥4 repeats, ≥60 chars), then trains LoRA with Final Token Preference Optimization (FTPO): one rejected token vs up to 20 chosen alternatives, logit-space KL, two-part regularization. Early LFM2.5-2.6B: 10.2%→1.4% doom-loop rate; Qwen3.5-4B greedy: 22.9%→1%, with eval gains attributed to fewer loops. Early-stop at chosen_win≈0.35. Prompt mix LiquidAI/antidoom-mix-v1.0 (478,229 rows) used to elicit loops.

Keywords

antidoom · ftpo · doom-loop · liquid-ai · lfm · qwen · repetition · blog · x

Topics

reasoning models, repetition, preference optimization

Research notes

  • Primary: Liquid AI blog (cite liquidAI2026Antidoom). Discord posted https://t.co/zPFEt8lMdr plus https://x.com/liquidai/status/2074494130126811473. Code https://github.com/Liquid4All/antidoom (Apache-2.0). Substantial prompt mix → datasets_local Antidoom Mix v1.0. Written by Sam Paech with listed contributors.