← Back to explorer

Language Models Are Implicitly Continuous

Type
other
Venue
arXiv / University of Oxford / University of Bologna / ETH Zurich

Summary

Defines a Continuous Causal Transformer (integral attention) that recovers discrete Transformers at unit duration. On Llama2/3, Phi-3, Gemma 1/2, and Mistral, shrinking token or sentence duration smoothly changes counts (about 190–205% more unique numeric peaks than a discrete model would allow). Linear interpolations of embeddings (apples–bananas) are treated as valid concepts (color, fruit-ness) even when they map to no vocabulary token. Shift-invariant, not scale-invariant. Published at ICLR 2025. Code https://github.com/samuelemarro/continuous-llm-experiments.

Keywords

continuous-llm · token-duration · embedding-interpolation · iclr · oxford · interpretability · cct

Topics

interpretability, continuous representations, Transformer theory

Research notes

  • Primary: arxiv abs (cs.CL; also cs.LG). License not stated on abs/HTML at check. Comment: Published at ICLR 2025. Oxford / Bologna / ETH Zurich. Correspondence samuele.marro@eng.ox.ac.uk. Code https://github.com/samuelemarro/continuous-llm-experiments (18 stars at check). HF paper page 3 upvotes; githubRepo linked; no linked models/datasets. Discord posted abs. Prompt experiments only; no new corpus, so no datasets_local row. License field left blank per catalog convention.