datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
deepresearchgym-agentic-search-logs
DeepResearchGym Agentic Search Logs
This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617).
The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253.
All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/deepresearchgym-agentic-search-logs.ASQ
🤖 ASQ: Agentic Search Queryset
A dataset capturing RAG agents' search behaviours.
📖 Dataset Description
ASQ (Agentic Search Queryset) is a dataset designed to capture the search behaviors of the RAG agents.
It collects intermediate synthetic queries, retrieved documents, and thoughts (reasoning descriptions) produced or consumed by agents.
📊 Dataset Statistics
615k traces (0.12% incomplete)
614k answers
680k synthetic queries
680k retrieved ranked… See the full description on the dataset page: https://huggingface.co/datasets/AgenticSearchQueryset/ASQ.Nemotron-SFT-Agentic-v2-search-toolcalling-parquet
Nemotron-SFT-Agentic-v2 Search and Tool Calling Parquet
Subset Parquet conversion of nvidia/Nemotron-SFT-Agentic-v2 containing only the search and tool calling splits.
Nested JSON fields are preserved as compact JSON strings to keep a stable Parquet schema across records.
Files
search.parquet: 5,968 rows
tool_calling.parquet: 8,444 rows
Note: one malformed source record in tool_calling.jsonl is preserved via __raw_record and __parse_error.
agentic-search-data
Agentic search — synthetic multi-hop retrieval dataset
JSONL artifacts for training and evaluating a retrieval agent across web, finance, legal, code, and science.
Files
File
Description
corpus.jsonl
Unique doc_id passages (supporting + distractor docs) with domain labels
sft_dataset.jsonl
Supervised fine-tuning tasks (~60% of tasks)
rl_dataset.jsonl
RL / GRPO-style prompts (~25%)
eval_dataset.jsonl
Held-out evaluation (~15%)
Splits are disjoint by… See the full description on the dataset page: https://huggingface.co/datasets/2796gauravc/agentic-search-data.repro-mm-deepresearch-a-simple-and-effective-multimodal-agentic-search-baseline-traces
Agent traces
Agent sessions published from a Trackio Logbook.
deepresearchgym-agentic-search-logs
DeepResearchGym Agentic Search Logs
This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617).
The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253.
All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/ethanning/deepresearchgym-agentic-search-logs.agentic-code-search-rollouts-sample10alpha_agentic_search_ragdataAgenticSearchQASearch-Agentic-v1
Search-Agentic-v1 Dataset
This dataset contains agentic trajectories (ReAct format) in English and Vietnamese designed for agentic fine-tuning. The agent has access to web search and web page reading tools.
Dataset Structure
Each record contains the following fields:
id: SHA-256 hash of the initial user prompt, serving as a unique identifier.
lang: Linguistic attribution, either "en" (English) or "vi" (Vietnamese).
tools: Comprehensive list of function tools… See the full description on the dataset page: https://huggingface.co/datasets/iselabvn/Search-Agentic-v1.agentic-search-rl-mixed-shortform-dr-tulu-longform-v1search_data_agenticAgentic-Search-Hardagentic-search-chromadbagentic-search-refiner-shortform-curated-v1
Agentic Search Refiner Shortform Curated v1
Curated short-form QA mixture for /home/nvidia/workspace/dr-tulu/rl/open-instruct/run_train_refiner_shortform_api_agent.sh.
This dataset is normalized to the ASearcher-compatible schema expected by rlvr_tokenize_asearcher_v1:
question: user query
answer: exact reference answer as a string
source, subset, source_id, metadata: provenance/debug fields
Composition
{
"asearcher_lrm_multihop": 300… See the full description on the dataset page: https://huggingface.co/datasets/lihaoxin2020/agentic-search-refiner-shortform-curated-v1.Agentic-Search-MultiHopagentic-search-dataagentic_dataset_search_v2agentic_dataset_searchAgentic-Search-SFT-V0Agentic-Search-PromptsAgentic-Search-SFT-SampleAgentic-Search-Hard-V0Agentic-Search-MultiHop-Easy-Set
