DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent
- Type
- dataset
- Venue
- City University of Hong Kong, Baidu, Shenzhen Technology University
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.740555+00:00
- Verified
- 2026-07-17T20:18:03.740555+00:00
Summary
DeepResearch-9K is a large-scale challenging dataset for deep-research agent training and evaluation, consisting of 9,000 questions spanning three difficulty levels (L1-L3) with high-quality search trajectories and reasoning chains generated by Tongyi-DeepResearch-30B-A3B. Built from open-source multi-hop QA datasets via a low-cost autonomous pipeline, it includes verifiable answers and a hard subset (DeepResearch-Hard, 3,974 samples where the teacher model failed). Used with the DeepResearch-R1 training framework for RL-based agent training.
Keywords
deep-research agent multi-hop-qa benchmark trajectory reinforcement-learning search reasoning
Topics
NLP / Deep Research Agents
Research notes
- Paper: arxiv 2603.01152, published at SIGIR 2026. Code at github.com/Applied-Machine-Learning-Lab/SIGIR2026_DeepResearch-R1. Train/test split designed so test set contains only hard/incorrect cases.