MANTA-1M
- Type
- dataset
- Venue
- LGAI-EXAONE
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en
- Added
- 2026-07-17T20:01:42.159104+00:00
- Verified
- 2026-07-17T20:01:42.159104+00:00
Summary
We introduce **MANTA**, an automated pipeline that generates high-quality large-scale instruction fine-tuning datasets from massive web corpora while preserving their diversity and scalability. By extracting structured syllabi from web documents and leveraging high-performance LLMs, our approach enables highly effective query-response generation with minimal human intervention. Extensive experiments on 8B-scale LLMs demonstrate that fine-tuning on the MANTA-1M dataset significantly outperforms other massive dataset generation methodologies, particularly in knowledge-intensive tasks such as MMLU and MMLU-Pro, while also delivering superior performance across a broad spectrum of tasks.
Keywords
hf-dataset qa parquet text datasets pandas polars mlcroissant has-paper
Topics
QA
Research notes
- downloads=164; likes=27