Comparing Transformers and Hybrid Models at the Token Level
- Type
- paper
- Venue
- arXiv / Allen Institute for AI
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:58:34Z
- Verified
- 2026-08-14T16:58:34Z
Summary
Paired NLL of matched Olmo 3 7B (transformer) and Olmo Hybrid 7B at identical prefixes. Hybrid advantage is broad but largest on open-class content words (0.0384 nats vs 0.0238 on function words) and on opening delimiters vs closers; it nearly vanishes on long repeated n-grams. Synthetic probes: hybrid wins pronoun-memory and entity-tracking; transformer wins structural closure (predicting the closer). Interprets this as recurrence helping discourse/program state readout while attention handles visible-prefix copy and bracket matching. Proof-of-concept 1B filtered evals (Top-10∩No-Copy vs Copy-5-only) roughly double the architecture gap vs aggregate validation. No official code on the abs page.
Keywords
olmo · hybrid · token-level · state-tracking · allenai · gdn
Topics
hybrid language models, interpretability, state tracking
Research notes
- Primary: arxiv abs (CC BY 4.0, cs.CL/AI). Allen Institute for AI. Uses released Olmo 3 7B and Olmo Hybrid 7B (Apache-2.0) and Olmo Hybrid paper 2604.03444; no new corpus, so no datasets_local row. HF has no paper page (404). No official code on abs. Discord posted abs link.