CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01muset-ai /DeepResearch-Bench-II-Datasetdocumenttext-generationn<1K2 likes1.4k downloads7mo agoHugging Face02muset-ai /DeepResearch-Bench-Dataset DeepResearch Bench Dataset [English | 中文] English 📖 Dataset Overview This is the official dataset accompanying the DeepResearch Bench paper. It contains research reports generated by 4 leading deep research AI systems along with detailed human expert annotations evaluating these reports. DeepResearch Bench is the first comprehensive benchmark for systematically evaluating Deep Research Agents (DRAs) on their ability to handle complex, PhD-level research… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/DeepResearch-Bench-Dataset.text-generationn<1K10 likes345 downloads10mo agoHugging Face03DanielTobi0 /openresearcher-sft-deep-research-cleaned OpenResearcher SFT DeepResearch — Parquet Mirror This is a re-hosted copy of the tool-reasoning SFT deep-research dataset by Aman Priyanshu, itself a cleaned/restructured version of the OpenResearcher Dataset from TIGER-AI-Lab. Why this repo exists: the source wasn't laid out as ready-to-download Parquet files. This mirror simply stores the data as plain seed_*.parquet files so you can grab the whole dataset or a single segment easily. No changes were made to the content — all… See the full description on the dataset page: https://huggingface.co/datasets/DanielTobi0/openresearcher-sft-deep-research-cleaned.tabulartext-generation10K<n<100K0 likes286 downloads2mo agoHugging Face04IPF /DeepResearch-traj DeepResearch-traj Multi-seed deep research agent trajectories with per-question correctness labels and pass@k statistics, derived from OpenResearcher/OpenResearcher-Dataset. Dataset Summary This dataset contains 97,630 full agent trajectories across 6,102 unique research questions, each sampled under 16 different random seeds (42–57). Every trajectory is annotated with: seed — which random seed produced this trajectory correct — whether the model's final answer was… See the full description on the dataset page: https://huggingface.co/datasets/IPF/DeepResearch-traj.tabularquestion-answering10K<n<100K0 likes277 downloads7mo agoHugging Face05AmanPriyanshu /tool-reasoning-sft-RESEARCH-openresearcher-dataset-sft-deep-research-agent-data-cleaned OpenResearcher Dataset - Cleaned & Restructured 👥 Follow the Author Aman Priyanshu Overview This dataset is a cleaned and restructured version of the OpenResearcher Dataset released by the TIGER-AI-Lab. The original dataset contains 96K+ long-horizon deep research trajectories generated by GPT-OSS-120B with native browser tools. This version converts the GPT-OSS channel-based message format into a standardized multi-turn tool-use conversation… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-openresearcher-dataset-sft-deep-research-agent-data-cleaned.text-generation10K<n<100K2 likes261 downloads6mo agoHugging Face06AlgorithmicResearchGroup /arxiv_deep_learning_python_research_code ArXiv Deep Learning Python Research Code A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code. Dataset Summary Statistic Value Total files 391,496 Total size 1.49 GB Source repos 34,099 Time span ArXiv inception through July 2023 Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code.tabulartext-generation100K<n<1M11 likes242 downloads6mo agoHugging Face07cx-cmu /deepresearchgym-agentic-search-logs DeepResearchGym Agentic Search Logs This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617). The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253. All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/deepresearchgym-agentic-search-logs.tabulartext-retrieval10M<n<100M16 likes222 downloads8mo agoHugging Face08InternScience /SGI-DeepResearchgated Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows &nbsp; &nbsp; &nbsp; Welcome to the official repository for the SGI-Bench! 👏 Scientist-aligned benchmark for evaluating Scientific General Intelligence (SGI) across the full inquiry cycle: Deliberation, Conception, Action, and Perception. The benchmark spans 10 disciplines and more than 1,000 expert‑curated samples inspired by Science’s 125 Big Questions, with an agentic evaluation framework… See the full description on the dataset page: https://huggingface.co/datasets/InternScience/SGI-DeepResearch.textquestion-answeringn<1K11 likes148 downloads4mo agoHugging Face09Alibaba-NLP /Open-DeepResearch Open-DeepResearch Project Page | Paper | Code This directory contains the RL Training Set and the Test Set for the Open-DeepResearch domain. Overview In the Open-DeepResearch domain, the agent is required to assist users in conducting multi-turn search, reading, synthesis, and generation to produce an open-ended answer. This domain focuses on complex information retrieval and synthesis tasks. Dataset Statistics Split Samples Description RL… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/Open-DeepResearch.text-generation6 likes147 downloads8mo agoHugging Face10JRQi /DeepResearch-Bench-Multilingual DeepResearch Bench Multilingual Prompts This dataset provides prompt-level multilingual translations for the 100 research tasks used in muset-ai/DeepResearch-Bench-Dataset. The translations cover eight languages: en zh es it ar bn ja el What is included This repository focuses on the benchmark prompts only. On the Hugging Face Hub, the Dataset Viewer is configured with one default subset named all plus nine explicit subset configurations: source_prompt, en, zh, es… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/DeepResearch-Bench-Multilingual.texttext-generation1K<n<10K1 likes113 downloads6mo agoHugging Face11xiesixiong /deepresearch-benchmark-2 DeepResearch Benchmark 2.0 DeepResearch Benchmark 2.0 is a collection of 100 English deep-research benchmark cases. Each case asks a model to analyze 6-10 entities across 6-10 research dimensions, and includes: the public user-facing question, a reference answer with derivations and source URLs, a detailed scoring rubric, metadata for the generation/auditing pipeline when available. This Hugging Face package is the clean OpenReview dataset release. It excludes local MCP configs… See the full description on the dataset page: https://huggingface.co/datasets/xiesixiong/deepresearch-benchmark-2.question-answering0 likes101 downloads5mo agoHugging Face12SupritiVijay /tool-reasoning-sft-RESEARCH-dr-tulu-sft-deep-research-agent-data-cleaned-rectified Deep Research - Tulu SFT Data Cleaned Rectified 👥 Follow the Author Supriti Vijay Overview This dataset is a cleaned and restructured version of the DR-TULU SFT dataset released by AllenAI's RL Research team. The original DR-TULU dataset represents significant work in creating high-quality training data for reasoning-enhanced language models with tool use capabilities. This version addresses structural issues in the original release while preserving… See the full description on the dataset page: https://huggingface.co/datasets/SupritiVijay/tool-reasoning-sft-RESEARCH-dr-tulu-sft-deep-research-agent-data-cleaned-rectified.tabulartext-generation10K<n<100K8 likes85 downloads10mo agoHugging Face13xsx001 /deepresearch-benchmark-2 DeepResearch Benchmark 2.0 DeepResearch Benchmark 2.0 is a collection of 100 English deep-research benchmark cases. Each case asks a model to analyze 6-10 entities across 6-10 research dimensions, and includes: the public user-facing question, a reference answer with derivations and source URLs, a detailed scoring rubric, metadata for the generation/auditing pipeline when available. This Hugging Face package is the clean OpenReview dataset release. It excludes local MCP configs… See the full description on the dataset page: https://huggingface.co/datasets/xsx001/deepresearch-benchmark-2.question-answering0 likes73 downloads5mo agoHugging Face14InternScience /SGI-DeepResearch-Goldgated Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows &nbsp; &nbsp; &nbsp; Welcome to the official repository for the SGI-Bench! 👏 Scientist-aligned benchmark for evaluating Scientific General Intelligence (SGI) across the full inquiry cycle: Deliberation, Conception, Action, and Perception. The benchmark spans 10 disciplines and more than 1,000 expert‑curated samples inspired by Science’s 125 Big Questions, with an agentic evaluation framework… See the full description on the dataset page: https://huggingface.co/datasets/InternScience/SGI-DeepResearch-Gold.textquestion-answeringn<1K4 likes55 downloads4mo agoHugging Face15ethanning /deepresearchgym-agentic-search-logs DeepResearchGym Agentic Search Logs This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617). The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253. All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/ethanning/deepresearchgym-agentic-search-logs.tabulartext-retrieval10M<n<100M1 likes33 downloads8mo agoHugging Face16tuandunghcmut /deepresearch-9k-tool-calling-strictgated DeepResearch-9K — Tool Calling Format (Strict) Strict converted version of artillerywu/DeepResearch-9K. Key difference from the standard version: When an assistant message contains tool_calls, the content field is null. <think> reasoning blocks are dropped from tool-calling turns. Dataset Summary Property Value Source artillerywu/DeepResearch-9K Samples 3,974 Tool search Format OpenAI-compatible messages + tools_json Difficulty… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/deepresearch-9k-tool-calling-strict.texttext-generation1K<n<10K0 likes15 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.