YourBench: Easy Custom Evaluation Sets for Everyone
- Type
- paper
- Venue
- arXiv / Hugging Face / UIUC
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T18:56:35Z
- Verified
- 2026-08-14T18:56:35Z
Summary
Open-source Document-to-Evaluation Generation pipeline (ingest PDF/Word/HTML to markdown, semantic chunk plus summary, ensemble QA generation with citations, fuzzy citation filter theta=0.85, SBERT+DBSCAN dedup). Replicates 7 MMLU subsets from a few Wikipedia pages for under $15 total / under $2 per domain and preserves model ranking (Spearman rho=1 on mean scores) while being harder. Human validity ~85% (2k questions, 20 annotators, Gwet AC1=0.71). Introduces Tempora-0325 (7,368 docs published after 2025-03-01) plus 150k+ Tempora QA pairs and inference traces. Evaluated 26 models (3B-671B, 7 families). COLM 2025. Code Apache-2.0 github.com/huggingface/yourbench. Posted HF Space yourbench/demo 404s at catalog time; GitHub still lists demo and advanced Spaces.
Keywords
yourbench · evaluation · synthetic-benchmarks · mmlu · tempora · huggingface · colm-2025 · d2eg
Topics
LLM evaluation, synthetic benchmarks
Research notes
- Primary: arxiv abs 2504.01833. Discord posted https://huggingface.co/spaces/yourbench/demo (404 at catalog; GitHub README still links demo + advanced Spaces). Code https://github.com/huggingface/yourbench (~450 stars, Apache-2.0). Site https://yourbench.github.io/. Tempora dataset https://huggingface.co/datasets/yourbench/tempora (also loadable as sumuks/tempora). Substantial source-doc corpus appended to datasets_local as Tempora-0325. Authors use accents (Clementine Fourrier; Dilek Hakkani-Tur) omitted here for CSV safety; names match the arXiv author list.