datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepResearch-Bench-II-Datasetdeep_research_bench_evaldeepresearch-corpusS1-DeepResearch-15k
S1-DeepResearch-15k Dataset
Overview
The S1-DeepResearch dataset is a curated collection of approximately 15k samples designed to improve deep research capabilities of large language models.
The dataset includes two types of tasks:
Verifiable tasks (labeled as "Closed-ended Multi-hop Resolution")
Open-ended tasks (labeled as "Open-ended Exploration")
Dataset Composition
The dataset is organized into five core capability dimensions:
Long-chain complex… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-DeepResearch-15k.AI-DeepResearch-BenchReportDeepResearch-Bench-Dataset
DeepResearch Bench Dataset
[English | 中文]
English
📖 Dataset Overview
This is the official dataset accompanying the DeepResearch Bench paper. It contains research reports generated by 4 leading deep research AI systems along with detailed human expert annotations evaluating these reports.
DeepResearch Bench is the first comprehensive benchmark for systematically evaluating Deep Research Agents (DRAs) on their ability to handle complex, PhD-level research… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/DeepResearch-Bench-Dataset.deepresearchMM-DeepResearch-corpusDeepResearch_wiki_data
📚 DeepResearch Wiki Data
A comprehensive collection of Wikipedia articles about recent events (2023-2025), stored in JSON format. This dataset serves as the source material for the CriticSearch Report Benchmark, containing detailed information about global events, elections, disasters, and other significant occurrences.
🌟 Dataset Overview
Time Period: 2023-2025
Total Articles: 271 articles
Format: JSON files
Source: Wikipedia
Last Updated: 2025-06-06
📊… See the full description on the dataset page: https://huggingface.co/datasets/Looogic/DeepResearch_wiki_data.hermes-bc-traj
Browse Comp Eval
Standalone parallel Hermes runner. It does not import AIDABench; it only follows the same operational shape: JSONL input, concurrent runs, retry, resume, and per-run JSON outputs.
Input
Put JSONL files under data/{dataset}/. Each row must contain:
{"question": "...", "answer": "...", "type": "..."}
answer is only recorded for later evaluation. It is not sent to Hermes.
Run
cd /root/Browse_comp_eval
export HERMES_API_KEY="..."… See the full description on the dataset page: https://huggingface.co/datasets/ICA-DeepResearch/hermes-bc-traj.DeepResearch-traj
DeepResearch-traj
Multi-seed deep research agent trajectories with per-question correctness labels and pass@k statistics, derived from OpenResearcher/OpenResearcher-Dataset.
Dataset Summary
This dataset contains 97,630 full agent trajectories across 6,102 unique research questions, each sampled under 16 different random seeds (42–57). Every trajectory is annotated with:
seed — which random seed produced this trajectory
correct — whether the model's final answer was… See the full description on the dataset page: https://huggingface.co/datasets/IPF/DeepResearch-traj.DeepResearch-9K
Data Splits
The dataset consists of the following two subsets:
Dataset Name
Contents
How It Was Generated
Number of Samples
DeepResearch-9K
All samples (teacher model's outputs)
Teacher model inference on 9K questions
9,000
DeepResearch-Hard
Teacher model's incorrect samples only
Filtered from DeepResearch-9K (samples where the teacher model's final answer was wrong)
3,974
Based on the above, we define the following train/test split for our DeepResearch-R1 model:… See the full description on the dataset page: https://huggingface.co/datasets/artillerywu/DeepResearch-9K.deepresearch-benchdeepresearchgym-agentic-search-logs
DeepResearchGym Agentic Search Logs
This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617).
The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253.
All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/deepresearchgym-agentic-search-logs.deepresearch-bench-queryOpen-DeepResearch
Open-DeepResearch
Project Page | Paper | Code
This directory contains the RL Training Set and the Test Set for the Open-DeepResearch domain.
Overview
In the Open-DeepResearch domain, the agent is required to assist users in conducting multi-turn search, reading, synthesis, and generation to produce an open-ended answer. This domain focuses on complex information retrieval and synthesis tasks.
Dataset
Statistics
Split
Samples
Description
RL… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/Open-DeepResearch.SGI-DeepResearch
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
Welcome to the official repository for the SGI-Bench! 👏
Scientist-aligned benchmark for evaluating Scientific General Intelligence (SGI) across the full inquiry cycle: Deliberation, Conception, Action, and Perception. The benchmark spans 10 disciplines and more than 1,000 expert‑curated samples inspired by Science’s 125 Big Questions, with an agentic evaluation framework… See the full description on the dataset page: https://huggingface.co/datasets/InternScience/SGI-DeepResearch.DeepResearch9K-qwen3-1p7b-eval-traces-0917
Qwen3-1.7B worker evaluation traces — DeepResearch9K
12 raw/SFT runs under six orchestrators. Published with user authorization to use public HF storage on 2026-09-17. Original private repositories and historical data remain unchanged; only this new 1.7B campaign is included here.
Full raw results, coordinator/worker traces, measured role tokens/cache, episode timings, input snapshot, manual-protocol grading, and separate theoretical within-conversation Max prefix estimates are… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/DeepResearch9K-qwen3-1p7b-eval-traces-0917.Deep-Research-Benchmarks
Deep Research Benchmarks
Password-protected bundle of the public deep-research benchmarks used by AgentHarness to evaluate Apodex-1.0 in standard ReAct mode.
Download
wget https://huggingface.co/datasets/apodex/Deep-Research-Benchmarks/resolve/main/deep_research_benchmarks_260607.zip
unzip -P 'apodex*()_2026' deep_research_benchmarks_260607.zip
rm deep_research_benchmarks_260607.zip
Single quotes around the password are required — it contains *, (, ).
After… See the full description on the dataset page: https://huggingface.co/datasets/apodex/Deep-Research-Benchmarks.DeepResearch-Bench-Multilingual
DeepResearch Bench Multilingual Prompts
This dataset provides prompt-level multilingual translations for the 100 research tasks used in muset-ai/DeepResearch-Bench-Dataset.
The translations cover eight languages:
en
zh
es
it
ar
bn
ja
el
What is included
This repository focuses on the benchmark prompts only.
On the Hugging Face Hub, the Dataset Viewer is configured with one default subset named all plus nine explicit subset configurations: source_prompt, en, zh, es… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/DeepResearch-Bench-Multilingual.deep_research_taskset_fulldeepresearch-v1deepresearch-benchmark-2
DeepResearch Benchmark 2.0
DeepResearch Benchmark 2.0 is a collection of 100 English deep-research benchmark cases.
Each case asks a model to analyze 6-10 entities across 6-10 research dimensions, and includes:
the public user-facing question,
a reference answer with derivations and source URLs,
a detailed scoring rubric,
metadata for the generation/auditing pipeline when available.
This Hugging Face package is the clean OpenReview dataset release. It excludes local MCP configs… See the full description on the dataset page: https://huggingface.co/datasets/xiesixiong/deepresearch-benchmark-2.Vision-DeepResearch-Evalepago-sn36-deepresearch-trajectories
Epago SN36 — trajectories, baselines, diagnostics and mining tooling
Everything produced while investigating model mining on Bittensor subnet 36 (Epago).
All data was generated locally by running the subnet's own harness against its own bundled
corpus. Nothing here is copied from the subnet's private artifacts.
⚠️ Read this first: coronation is currently impossible
On EpagoFoundation/epago @ 7ddfef0 (latest origin/main as of 2026-09-10), no challenger
can ever be… See the full description on the dataset page: https://huggingface.co/datasets/Olague-Secret/epago-sn36-deepresearch-trajectories.Pre-Training-Persian-Corpus-Raw-Texts-DatasetDocument Version: 2.0.0 | Last Updated: 02/13/2026
DeepResearch-SFT
Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval And Synthesis For SLMs
✨ News
[29/09/25]: Our paper on Fathom-Search-4B has been accepted to SEA @ NeurIPS 2025 🎉 OpenReview link
Introduction
We introduce Fathom-DeepResearch, an agentic DeepResearch system that sets state-of-the-art performance in the open-weights category on search-intensive benchmarks (SimpleQA, FRAMES, WebWalkerQA, Seal0) and outperforms… See the full description on the dataset page: https://huggingface.co/datasets/FractalAIResearch/DeepResearch-SFT.deer_sample_queriesDeepResearch-Datadeep_research_taskset
