← Back to explorer

stream-data

Type
dataset
Venue
JonasGeiping / Hugging Face
Year
2026
Source
huggingface
Access
free
Language
English
Added
2026-07-17T20:18:35.164762+00:00
Verified
2026-07-17T20:18:35.164762+00:00

Summary

A training corpus for the Stream-LLM models (Stream-Qwen3.5-27B, Stream-Qwen3-8B) that contains 3,874 machine-generated samples in a 10-column grid format, where each column represents one cognitive channel (User, Output, Analytical, Skeptical, Intuitive, Between, Curious, Void, Instinct, Synthesis) and each row is one timestep. Streams were synthesized via the Anthropic API (Claude Opus 4.5) given an input prompt and a system message describing the ten-channel protocol. The dataset supports research on multi-stream LLMs that unblock language models by splitting roles into separate parallel streams of computation, enabling simultaneous reading, thinking, and output generation in a single forward pass.

Keywords

stream-llm multi-stream parallel-cognition synthesized reasoning cognitive-channels qwen

Topics

NLP / LLM Reasoning

Research notes

  • Paper: 'Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs' (arxiv 2605.12460). Code at github.com/seal-rg/streaming. The processed config is tokenized with the Qwen3.5-27B tokenizer (vocab 248320, silence token id 481). Block-causal attention prevents same-row tokens from seeing each other.