paper_instructions_300K-v1
- Type
- dataset
- Venue
- Paper Breakdown
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.795215+00:00
- Verified
- 2026-07-17T20:18:03.795215+00:00
Summary
paper_instructions_300K-v1 contains ~300K synthetic Alpaca-style instruction-response pairs generated from 1,500 ML/AI papers using text-albumentations. It transforms long-form technical text into diverse, task-shaped supervision covering summarization, question generation, fact extraction, and reasoning tasks. The dataset has been used to fine-tune SmolLM-135M into a structured research assistant API for ML papers.
Keywords
instruction-tuning synthetic papers sft distillation ml arxiv alpaca
Topics
NLP / ML papers
Research notes
- Generated via text-albumentations chunking and constrained synthetic generation pipeline. Includes a reasoning subset (~11.8k rows) added later. Train/test split: 254k/22.2k.