datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthesized_datasetS1-DeepResearch-15k
S1-DeepResearch-15k Dataset
Overview
The S1-DeepResearch dataset is a curated collection of approximately 15k samples designed to improve deep research capabilities of large language models.
The dataset includes two types of tasks:
Verifiable tasks (labeled as "Closed-ended Multi-hop Resolution")
Open-ended tasks (labeled as "Open-ended Exploration")
Dataset Composition
The dataset is organized into five core capability dimensions:
Long-chain complex… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-DeepResearch-15k.simple-evalsdeepresearch-bench-queryDeepResearch-SFT
Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval And Synthesis For SLMs
✨ News
[29/09/25]: Our paper on Fathom-Search-4B has been accepted to SEA @ NeurIPS 2025 🎉 OpenReview link
Introduction
We introduce Fathom-DeepResearch, an agentic DeepResearch system that sets state-of-the-art performance in the open-weights category on search-intensive benchmarks (SimpleQA, FRAMES, WebWalkerQA, Seal0) and outperforms… See the full description on the dataset page: https://huggingface.co/datasets/FractalAIResearch/DeepResearch-SFT.Vision-DeepResearch-Text-DataDeepResearchEvalThis dataset contains 100 high-quality deep research tasks from DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation.
GitHub repository: https://github.com/Infinity-AILab/DeepResearchEval
repro-mm-deepresearch-a-simple-and-effective-multimodal-agentic-search-baseline-traces
Agent traces
Agent sessions published from a Trackio Logbook.
ICA-SFT-14kdeepresearch-bench-reference-cleandeepresearch-bench-criteriabiomedicine-deep-research
Biomedicine Deep Research
Complete materialized data for the 13-benchmark biomedicine deep research track. The files preserve the host materializer's directory layout.
public/development/: labeled fit and tune examples.
public/verifier/: unlabeled evaluation inputs.
public/reference/: audited biomedical source allowlist, RiskCalcs, and its notice.
private/: evaluation labels, mounted only into the separate grader during benchmark runs.
The complete tree is downloadable from… See the full description on the dataset page: https://huggingface.co/datasets/zifeng-ai/biomedicine-deep-research.ICA-RL-Sample-400deep-research-agent
Deep Research Agent Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/deep-research-agent.deepresearch-bench-reference-rawsynthetic_deep_research_dataset_gpt4ominiSam-R-deep-researchDeepResearchPrivate
