CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01muset-ai /DeepResearch-Bench-Dataset DeepResearch Bench Dataset [English | 中文] English 📖 Dataset Overview This is the official dataset accompanying the DeepResearch Bench paper. It contains research reports generated by 4 leading deep research AI systems along with detailed human expert annotations evaluating these reports. DeepResearch Bench is the first comprehensive benchmark for systematically evaluating Deep Research Agents (DRAs) on their ability to handle complex, PhD-level research… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/DeepResearch-Bench-Dataset.text-generationn<1K10 likes377 downloads10mo agoHugging Face02IPF /DeepResearch-traj DeepResearch-traj Multi-seed deep research agent trajectories with per-question correctness labels and pass@k statistics, derived from OpenResearcher/OpenResearcher-Dataset. Dataset Summary This dataset contains 97,630 full agent trajectories across 6,102 unique research questions, each sampled under 16 different random seeds (42–57). Every trajectory is annotated with: seed — which random seed produced this trajectory correct — whether the model's final answer was… See the full description on the dataset page: https://huggingface.co/datasets/IPF/DeepResearch-traj.tabularquestion-answering10K<n<100K0 likes287 downloads7mo agoHugging Face03InternScience /SGI-DeepResearchgated Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows &nbsp; &nbsp; &nbsp; Welcome to the official repository for the SGI-Bench! 👏 Scientist-aligned benchmark for evaluating Scientific General Intelligence (SGI) across the full inquiry cycle: Deliberation, Conception, Action, and Perception. The benchmark spans 10 disciplines and more than 1,000 expert‑curated samples inspired by Science’s 125 Big Questions, with an agentic evaluation framework… See the full description on the dataset page: https://huggingface.co/datasets/InternScience/SGI-DeepResearch.textquestion-answeringn<1K11 likes146 downloads4mo agoHugging Face04CharlieLLL /DeepResearch9K-qwen3-1p7b-eval-traces-0917 Qwen3-1.7B worker evaluation traces — DeepResearch9K 12 raw/SFT runs under six orchestrators. Published with user authorization to use public HF storage on 2026-09-17. Original private repositories and historical data remain unchanged; only this new 1.7B campaign is included here. Full raw results, coordinator/worker traces, measured role tokens/cache, episode timings, input snapshot, manual-protocol grading, and separate theoretical within-conversation Max prefix estimates are… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/DeepResearch9K-qwen3-1p7b-eval-traces-0917.question-answering0 likes144 downloads7d agoHugging Face05apodex /Deep-Research-Benchmarks Deep Research Benchmarks Password-protected bundle of the public deep-research benchmarks used by AgentHarness to evaluate Apodex-1.0 in standard ReAct mode. Download wget https://huggingface.co/datasets/apodex/Deep-Research-Benchmarks/resolve/main/deep_research_benchmarks_260607.zip unzip -P 'apodex*()_2026' deep_research_benchmarks_260607.zip rm deep_research_benchmarks_260607.zip Single quotes around the password are required — it contains *, (, ). After… See the full description on the dataset page: https://huggingface.co/datasets/apodex/Deep-Research-Benchmarks.question-answering3 likes123 downloads4mo agoHugging Face06JRQi /DeepResearch-Bench-Multilingual DeepResearch Bench Multilingual Prompts This dataset provides prompt-level multilingual translations for the 100 research tasks used in muset-ai/DeepResearch-Bench-Dataset. The translations cover eight languages: en zh es it ar bn ja el What is included This repository focuses on the benchmark prompts only. On the Hugging Face Hub, the Dataset Viewer is configured with one default subset named all plus nine explicit subset configurations: source_prompt, en, zh, es… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/DeepResearch-Bench-Multilingual.texttext-generation1K<n<10K1 likes120 downloads6mo agoHugging Face07xiesixiong /deepresearch-benchmark-2 DeepResearch Benchmark 2.0 DeepResearch Benchmark 2.0 is a collection of 100 English deep-research benchmark cases. Each case asks a model to analyze 6-10 entities across 6-10 research dimensions, and includes: the public user-facing question, a reference answer with derivations and source URLs, a detailed scoring rubric, metadata for the generation/auditing pipeline when available. This Hugging Face package is the clean OpenReview dataset release. It excludes local MCP configs… See the full description on the dataset page: https://huggingface.co/datasets/xiesixiong/deepresearch-benchmark-2.question-answering0 likes106 downloads5mo agoHugging Face08Olague-Secret /epago-sn36-deepresearch-trajectories Epago SN36 — trajectories, baselines, diagnostics and mining tooling Everything produced while investigating model mining on Bittensor subnet 36 (Epago). All data was generated locally by running the subnet's own harness against its own bundled corpus. Nothing here is copied from the subnet's private artifacts. ⚠️ Read this first: coronation is currently impossible On EpagoFoundation/epago @ 7ddfef0 (latest origin/main as of 2026-09-10), no challenger can ever be… See the full description on the dataset page: https://huggingface.co/datasets/Olague-Secret/epago-sn36-deepresearch-trajectories.question-answeringn<1K0 likes104 downloads14d agoHugging Face09xsx001 /deepresearch-benchmark-2 DeepResearch Benchmark 2.0 DeepResearch Benchmark 2.0 is a collection of 100 English deep-research benchmark cases. Each case asks a model to analyze 6-10 entities across 6-10 research dimensions, and includes: the public user-facing question, a reference answer with derivations and source URLs, a detailed scoring rubric, metadata for the generation/auditing pipeline when available. This Hugging Face package is the clean OpenReview dataset release. It excludes local MCP configs… See the full description on the dataset page: https://huggingface.co/datasets/xsx001/deepresearch-benchmark-2.question-answering0 likes77 downloads5mo agoHugging Face10InternScience /SGI-DeepResearch-Goldgated Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows &nbsp; &nbsp; &nbsp; Welcome to the official repository for the SGI-Bench! 👏 Scientist-aligned benchmark for evaluating Scientific General Intelligence (SGI) across the full inquiry cycle: Deliberation, Conception, Action, and Perception. The benchmark spans 10 disciplines and more than 1,000 expert‑curated samples inspired by Science’s 125 Big Questions, with an agentic evaluation framework… See the full description on the dataset page: https://huggingface.co/datasets/InternScience/SGI-DeepResearch-Gold.textquestion-answeringn<1K4 likes50 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.