← Back to explorer

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

Type
other
Venue
arXiv / OneLineAI / EleutherAI / CMU / Seoul National University

Summary

SOOHAK Challenge (340) and Refusal (99) plus companion SOOHAK-Mini (702), newly written by 105 mathematicians under NDA/IP transfer (~$550k MSIT Sovereign-AI budget). Challenge: Gemini-3-Pro / GPT-5 / Claude-Opus-4.5 Avg@3 30.4/26.4/10.4%; best open-weight Kimi-2.5 13.9%; 124 Challenge items unsolved by any of 11 models. Refusal (diagnose ill-posed prompts): no model >50% Avg@3, GLM-5 leads 49.5%. Mini: GPT-5 72.2%. Human baseline 5 teams / 25 solvers cover 50.6% of a 79-item slice; only Gemini-3-Pro exceeds combined humans. Public release planned late 2026; evals on request via guijin.son@snu.ac.kr. Project https://novamath.github.io.

Keywords

soohak · math-benchmark · research-math · refusal · contamination · imo · frontiermath · eleutherai · cmu · snu

Topics

math reasoning, benchmarks, contamination

Research notes

  • Primary: arxiv abs (cs.CL). License not stated on abs/HTML at check. Organizing team listed as authors (OneLineAI/EleutherAI/CMU/SNU and partners); 64+ mathematician dataset contributors omitted from authors field. Correspondence guijin.son@snu.ac.kr. No official code or public dataset on abs — full collection embargoed until late 2026 (NeurIPS 2026 acceptance window); no datasets_local row. HF paper page 82 upvotes; no linked models/datasets. Discord posted HF papers URL. Sample/eval-request workflow described; bilingual EN/KO planned.