← Back to explorer

CHIMERA

Type
dataset
Venue
TianHongZXY
Year
2026
Source
huggingface
Access
free
Language
en
Added
2026-07-17T20:00:46.239774+00:00
Verified
2026-07-17T20:00:46.239774+00:00

Summary

CHIMERA is a **compact but high-difficulty synthetic reasoning dataset** with **long Chain-of-Thought (CoT) trajectories** and **broad STEM coverage**, designed for **reasoning post-training**. All examples are **fully LLM-generated** and **automatically verified** without human annotation.

Keywords

hf-dataset language-modeling · -qa · -reasoning machine-generated parquet optimized-parquet text datasets dask polars mlcroissant reasoning chain-of-thought synthetic-data llm stem post-training has-paper

Topics

Language Modeling, QA, Reasoning

Research notes

  • downloads=292; likes=22