datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepResearch-traj
DeepResearch-traj
Multi-seed deep research agent trajectories with per-question correctness labels and pass@k statistics, derived from OpenResearcher/OpenResearcher-Dataset.
Dataset Summary
This dataset contains 97,630 full agent trajectories across 6,102 unique research questions, each sampled under 16 different random seeds (42–57). Every trajectory is annotated with:
seed — which random seed produced this trajectory
correct — whether the model's final answer was… See the full description on the dataset page: https://huggingface.co/datasets/IPF/DeepResearch-traj.deepresearchgym-agentic-search-logs
DeepResearchGym Agentic Search Logs
This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617).
The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253.
All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/deepresearchgym-agentic-search-logs.SGI-DeepResearch
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
Welcome to the official repository for the SGI-Bench! 👏
Scientist-aligned benchmark for evaluating Scientific General Intelligence (SGI) across the full inquiry cycle: Deliberation, Conception, Action, and Perception. The benchmark spans 10 disciplines and more than 1,000 expert‑curated samples inspired by Science’s 125 Big Questions, with an agentic evaluation framework… See the full description on the dataset page: https://huggingface.co/datasets/InternScience/SGI-DeepResearch.DeepResearch-Bench-Multilingual
DeepResearch Bench Multilingual Prompts
This dataset provides prompt-level multilingual translations for the 100 research tasks used in muset-ai/DeepResearch-Bench-Dataset.
The translations cover eight languages:
en
zh
es
it
ar
bn
ja
el
What is included
This repository focuses on the benchmark prompts only.
On the Hugging Face Hub, the Dataset Viewer is configured with one default subset named all plus nine explicit subset configurations: source_prompt, en, zh, es… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/DeepResearch-Bench-Multilingual.SGI-DeepResearch-Gold
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
Welcome to the official repository for the SGI-Bench! 👏
Scientist-aligned benchmark for evaluating Scientific General Intelligence (SGI) across the full inquiry cycle: Deliberation, Conception, Action, and Perception. The benchmark spans 10 disciplines and more than 1,000 expert‑curated samples inspired by Science’s 125 Big Questions, with an agentic evaluation framework… See the full description on the dataset page: https://huggingface.co/datasets/InternScience/SGI-DeepResearch-Gold.deepresearchgym-agentic-search-logs
DeepResearchGym Agentic Search Logs
This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617).
The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253.
All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/ethanning/deepresearchgym-agentic-search-logs.deepresearch-9k-tool-calling-strict
DeepResearch-9K — Tool Calling Format (Strict)
Strict converted version of artillerywu/DeepResearch-9K.
Key difference from the standard version:
When an assistant message contains tool_calls, the content field is null.
<think> reasoning blocks are dropped from tool-calling turns.
Dataset Summary
Property
Value
Source
artillerywu/DeepResearch-9K
Samples
3,974
Tool
search
Format
OpenAI-compatible messages + tools_json
Difficulty… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/deepresearch-9k-tool-calling-strict.
