← Back to explorer

DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent

Type
dataset
Venue
City University of Hong Kong, Baidu, Shenzhen Technology University
Year
2026
Source
huggingface
Access
free
Language
English
Added
2026-07-17T20:18:03.740555+00:00
Verified
2026-07-17T20:18:03.740555+00:00

Summary

DeepResearch-9K is a large-scale challenging dataset for deep-research agent training and evaluation, consisting of 9,000 questions spanning three difficulty levels (L1-L3) with high-quality search trajectories and reasoning chains generated by Tongyi-DeepResearch-30B-A3B. Built from open-source multi-hop QA datasets via a low-cost autonomous pipeline, it includes verifiable answers and a hard subset (DeepResearch-Hard, 3,974 samples where the teacher model failed). Used with the DeepResearch-R1 training framework for RL-based agent training.

Keywords

deep-research agent multi-hop-qa benchmark trajectory reinforcement-learning search reasoning

Topics

NLP / Deep Research Agents

Research notes

  • Paper: arxiv 2603.01152, published at SIGIR 2026. Code at github.com/Applied-Machine-Learning-Lab/SIGIR2026_DeepResearch-R1. Train/test split designed so test set contains only hard/incorrect cases.