← Back to explorer

TAPS-Datasets

Type
dataset
Venue
zbeeb (individual researcher)
Year
2026
Source
huggingface
Access
free
Language
English
Added
2026-07-17T20:18:03.793312+00:00
Verified
2026-07-17T20:18:03.793312+00:00

Summary

TAPS-Datasets contains 210k rows across three splits (MathInstruct, ShareGPT, and mixed) used to train lightweight draft models for speculative decoding. The accompanying paper 'TAPS: Task Aware Proposal Distributions for Speculative Sampling' (arXiv:2603.27027) studies how draft training data distribution affects speculative decoding quality, finding that task-specific training yields clear specialization: MathInstruct-trained drafts excel at reasoning benchmarks while ShareGPT-trained drafts excel at MT-Bench.

Keywords

speculative-decoding math chat draft-model inference-acceleration mathinstruct sharegpt

Topics

NLP / Speculative decoding

Research notes

  • Associated with GitHub repo Moe-Zbeeb/TAPS. Uses Meta-Llama-3-8B-Instruct as verifier and HASS/EAGLE-2 (~0.8B params) as drafters. Paper published March 2026.