TAPS-Datasets
- Type
- dataset
- Venue
- zbeeb (individual researcher)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.793312+00:00
- Verified
- 2026-07-17T20:18:03.793312+00:00
Summary
TAPS-Datasets contains 210k rows across three splits (MathInstruct, ShareGPT, and mixed) used to train lightweight draft models for speculative decoding. The accompanying paper 'TAPS: Task Aware Proposal Distributions for Speculative Sampling' (arXiv:2603.27027) studies how draft training data distribution affects speculative decoding quality, finding that task-specific training yields clear specialization: MathInstruct-trained drafts excel at reasoning benchmarks while ShareGPT-trained drafts excel at MT-Bench.
Keywords
speculative-decoding math chat draft-model inference-acceleration mathinstruct sharegpt
Topics
NLP / Speculative decoding
Research notes
- Associated with GitHub repo Moe-Zbeeb/TAPS. Uses Meta-Llama-3-8B-Instruct as verifier and HASS/EAGLE-2 (~0.8B params) as drafters. Paper published March 2026.