Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
- Type
- paper
- Venue
- arXiv / MPI-IS / ELLIS Tübingen
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T20:26:00Z
- Verified
- 2026-08-14T20:26:00Z
Summary
MPI-IS/Tübingen/ETH (cite Su et al.; arXiv 2605.12460). Instruction-tune for H parallel streams with intra-stream and cross-stream causality (attend to all other streams at t'<t). Interleaved packing + per-stream RoPE/stream embeddings; empty "-" slots masked with no KV. Synthetic wait-k / tabular stream data with causal verification. On Qwen3-1.7B/4B, overlapping read+solve drops TNFT to 0 and cuts delay while preserving GSM8K/MATH/LogicNLI/SQuAD accuracy; adding an audit stream beats sequential reflection at lower MSL. Stream isolation lowers prompt-injection ASR (StruQ-ID −33+ pp) without adversarial training. Extra internal streams raise eval-awareness/sub-vocalization and monitor accuracy vs same-family CoT (Qwen3.5-27B 10-stream). Small SFT vs production post-training; dense cross-stream attention.
Keywords
multi-stream · parallel-decoding · instruction-tuning · prompt-injection · monitorability · geiping
Topics
instruction tuning, parallel decoding, agent interfaces
Research notes
- Primary: arxiv abs 2605.12460 (cs.LG). Code https://github.com/seal-rg/streaming (no SPDX). Discord also linked Jonas Geiping X status. Not a hosted corpus.