← Back to explorer

Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs

Type
paper
Venue
arXiv / MPI-IS / ELLIS Tübingen
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T20:26:00Z
Verified
2026-08-14T20:26:00Z

Summary

MPI-IS/Tübingen/ETH (cite Su et al.; arXiv 2605.12460). Instruction-tune for H parallel streams with intra-stream and cross-stream causality (attend to all other streams at t'<t). Interleaved packing + per-stream RoPE/stream embeddings; empty "-" slots masked with no KV. Synthetic wait-k / tabular stream data with causal verification. On Qwen3-1.7B/4B, overlapping read+solve drops TNFT to 0 and cuts delay while preserving GSM8K/MATH/LogicNLI/SQuAD accuracy; adding an audit stream beats sequential reflection at lower MSL. Stream isolation lowers prompt-injection ASR (StruQ-ID −33+ pp) without adversarial training. Extra internal streams raise eval-awareness/sub-vocalization and monitor accuracy vs same-family CoT (Qwen3.5-27B 10-stream). Small SFT vs production post-training; dense cross-stream attention.

Keywords

multi-stream · parallel-decoding · instruction-tuning · prompt-injection · monitorability · geiping

Topics

instruction tuning, parallel decoding, agent interfaces

Research notes

  • Primary: arxiv abs 2605.12460 (cs.LG). Code https://github.com/seal-rg/streaming (no SPDX). Discord also linked Jonas Geiping X status. Not a hosted corpus.