← Back to explorer

paper_instructions_300K-v1

Type
dataset
Venue
Paper Breakdown
Year
2026
Source
huggingface
Access
free
Language
English
Added
2026-07-17T20:18:03.795215+00:00
Verified
2026-07-17T20:18:03.795215+00:00

Summary

paper_instructions_300K-v1 contains ~300K synthetic Alpaca-style instruction-response pairs generated from 1,500 ML/AI papers using text-albumentations. It transforms long-form technical text into diverse, task-shaped supervision covering summarization, question generation, fact extraction, and reasoning tasks. The dataset has been used to fine-tune SmolLM-135M into a structured research assistant API for ML papers.

Keywords

instruction-tuning synthetic papers sft distillation ml arxiv alpaca

Topics

NLP / ML papers

Research notes

  • Generated via text-albumentations chunking and constrained synthetic generation pipeline. Includes a reasoning subset (~11.8k rows) added later. Train/test split: 254k/22.2k.