CHIMERA
- Type
- dataset
- Venue
- TianHongZXY
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en
- Added
- 2026-07-17T20:00:46.239774+00:00
- Verified
- 2026-07-17T20:00:46.239774+00:00
Summary
CHIMERA is a **compact but high-difficulty synthetic reasoning dataset** with **long Chain-of-Thought (CoT) trajectories** and **broad STEM coverage**, designed for **reasoning post-training**. All examples are **fully LLM-generated** and **automatically verified** without human annotation.
Keywords
hf-dataset language-modeling · -qa · -reasoning machine-generated parquet optimized-parquet text datasets dask polars mlcroissant reasoning chain-of-thought synthetic-data llm stem post-training has-paper
Topics
Language Modeling, QA, Reasoning
Research notes
- downloads=292; likes=22