datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ResearchClawBench
ResearchClawBench
Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery
Quick Start | Submit Tasks | How It Works | Domains | Leaderboard | Add Your Agent
ResearchClawBench is a benchmark that measures whether AI coding agents can independently conduct scientific research — from reading raw data to producing publication-quality reports — and then rigorously evaluates the results against real human-authored papers.… See the full description on the dataset page: https://huggingface.co/datasets/InternScience/ResearchClawBench.newswire
Dataset Card for NewsWire
Dataset Summary
NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model.
Languages
English (en)
Dataset Structure
Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.lca-results
Long Code Arena (raw results)
These are the raw results from the Long Code Arena benchmark suite, as well as the corresponding model predictions.
Please use the subset dropdown menu to select the necessary data relating to our six benchmarks:
🤗 Library-based code generation
🤗 CI builds repair
🤗 Project-level code completion
🤗 Commit message generation🤗 Bug localization
🤗 Module summarization
data-product-benchmark
DPDisc Dataset
Paper | Code
Dataset Description
This dataset provides a benchmark for automatic data product creation. The task is framed as follows: given a natural language data product request and a corpus of text and tables, the objective is to identify the relevant tables and text documents that should be included in the resulting data product which would useful to the given data product request. The benchmark brings together three variants: HybridQA, TAT-QA, and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/data-product-benchmark.researchscope-papers
ResearchScope Papers
Open CS research paper dataset maintained by ResearchScope.
Updated automatically via GitHub Actions.
Quick start
from datasets import load_dataset
ds = load_dataset("kishormorol/researchscope-papers", "papers", split="train")
print(ds[0])
See Usage below for per-source splits, instruction-tuning, and the per-section fine-tuning data.
Stats
34,906 papers (raw metadata) — 9,906 arXiv · 20,000 conference · 5,000 journal
174,082… See the full description on the dataset page: https://huggingface.co/datasets/kishormorol/researchscope-papers.EnvBench
🌱⚙️ EnvBench
This repository contains data associated with EnvBench benchmark from EnvBench: A Benchmark for Automated Environment Setup.
It contains:
statistics about repositories from GitHub Search under ghs/data folder;
several data splits under splits folder;
Git repositories under repos folder;
READMEs under readmes folder (available as readmes config);
GitHub Actions workflows under workflows folder (available as workflows config);
list of files in the repositories… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/EnvBench.llama-9b-bulk-npzDREAM-1K
DREAM-1K
DREAM-1K (Description
with Rich Events, Actions, and Motions) is a challenging video description benchmark. It contains a collection of 1,000 short (around 10 seconds) video clips with diverse complexities from five different origins: live-action movies, animated movies, stock videos, long YouTube videos, and TikTok-style short videos. We provide a fine-grained manual annotation for each video.
Bellow is the dataset statistics:
Our1-2b-Datasetnoteflow-research-pilots
Keep the failed attempts. Check the artifact.
Versioned public development evidence from Robot Reel × Skills Anywhere × EvalArc, recorded 14 September 2026 on an NVIDIA L40S, with separate scripted Harbor controls on CPU and separate GPU context-control and agent-requested MCP handoff cohorts recorded 19 September 2026. This is an inspectable engineering casebook, not a held-out benchmark or training corpus with established efficacy.
Configuration
Actual experiment
What… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/noteflow-research-pilots.researchy_questions
Introduction
Researchy Questions is a set of about 100k Bing queries that users spent the most effort on. After a labor-intensive filtering funnel from billions of queries, these "needles in the haystack" are non-factoid, multi-perspective questions that probably require a lot of sub-questions and research in order to answer adequetly. These questions are shown to be harder than other open domain QA datasets like Natural Questions.
The train dataset has about 90k samples.… See the full description on the dataset page: https://huggingface.co/datasets/corbyrosset/researchy_questions.dr-tulu-sft-data
[!NOTE]
For full information, go check out the Dr Tulu paper here.
DR Tulu SFT Data
This dataset contains the SFT training data for DR Tulu, containing prompts and full trajectories including reasoning traces, tool calls, and answers with citations.
The source prompts are curated from OpenScholar, Search Arena, and short-form QA datasets inclduing WebWalker-Silver, TaskCraft, PopQA and TyDiQA (English).
Important: This does not contain the SFT subsets created using prompts from… See the full description on the dataset page: https://huggingface.co/datasets/rl-research/dr-tulu-sft-data.OR-Clarify
OR-Clarify
📄 Paper: Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
OR-Clarify is a benchmark for testing whether an agent asks the right questions before formulating an optimization model.
Most optimization benchmarks give an agent a complete problem statement. OR-Clarify instead starts with an incomplete business brief. The agent must identify missing requirements that could change the optimization formulation, ask for the relevant… See the full description on the dataset page: https://huggingface.co/datasets/AIOR-Research/OR-Clarify.ouroboros-pasterski-research-program
Ouroboros Pasterski Research Program - PAUSED
A proof-carrying research campaign in which the Ouroboros AI System turned public scholarship into 49 exact, independently replayable computational result packages. The campaign ran for 18 h 11 min before the operator paused it at a stable boundary; PAUSED means the research program remains open, not that the released results are provisional.
Independent project; no affiliation or endorsement. This work is not affiliated with… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-pasterski-research-program.verified-research-reasoning-trajectories
Verified Research Reasoning Trajectories for RLVR
This repository is the public sample and schema repository for Ulam's research-level mathematical reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), process supervision, judge training, proof criticism, and private evaluations.
Ulam Verified Research Reasoning Trajectories are proof-process data for RLVR. Each record contains a normalized research problem, a golden or partial-golden proof graph… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-research-reasoning-trajectories.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/vpasx/lgg-mri-segmentation-research.IndustryInstruction_Technology-Research
IndustryInstruction: Technology & Research
This repository contains the IndustryInstruction: Technology & Research domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Technology-Research.model-toolcall-research
Model Toolcall Research
kimicode_swetogether_traces
kimi-code × SWE-Together agentic traces
This is a dataset generated by a real agentic coding workload: SWE-Together tasks executed by the kimi-code agent, with every LLM call captured at the wire level. It is intended primarily as an inference-serving workload (in the spirit of Inferact/codex_swebenchpro_traces): full multi-turn sessions preserve the request structure — growing contexts, shared prefixes, tool calls — that drives KV-cache behavior in production serving.… See the full description on the dataset page: https://huggingface.co/datasets/verda-research/kimicode_swetogether_traces.ResearchClawBench
ResearchClawBench
Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery
Quick Start | Submit Tasks | How It Works | Domains | Leaderboard | Add Your Agent
ResearchClawBench is a benchmark that measures whether AI coding agents can independently conduct scientific research — from reading raw data to producing publication-quality reports — and then rigorously evaluates the results against real human-authored papers.… See the full description on the dataset page: https://huggingface.co/datasets/Ka3de/ResearchClawBench.r9-research-framework
R9 Research Framework — Qwen3.5-9B Distillation
⚠️ CRITICAL: READ FIRST — Ollama Inference Flag Required
If you serve any Qwen3.5-derived model from this lineage via Ollama,
you MUST pass "think": false in the /api/chat request body.
curl -X POST http://localhost:11434/api/chat \
-d '{"model": "qwen3.5-9b-r10:q4km", "think": false, "messages": [...], "stream": false}'
Without this flag the model will appear to "loop" and produce empty answers
on 25-46% of requests.… See the full description on the dataset page: https://huggingface.co/datasets/cudabenchmarktest/r9-research-framework.federal-register-live-graphrag-research-20260810
Federal Register live GraphRAG (research)
Local LCR-071 live pipeline output for the 2026-08-10 cutoff (11,784 documents,
CUDA thenlper/gte-small). This Hub copy is a research snapshot.
It is not a current-bundle and does not replace
justicedao/ipfs_federal_register. LCR-084 remains open. Official Federal
Register publications remain the authority.
Hub git directories may contain at most 10,000 files. Document bodies beyond
that cap are stored under corpus/bodies-part2/ rather… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/federal-register-live-graphrag-research-20260810.FrescoArchive
FrescoArchive
FrescoArchive is an evaluation dataset for large-format image-to-video generation. It contains 371 complex, multi-scene artworks paired with detailed English prompts. The collection was introduced with FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion.
This repository publishes provenance metadata and source links only. It does not redistribute the image files.
Dataset structure
Each row contains:
frescoarchive_id: the… See the full description on the dataset page: https://huggingface.co/datasets/obvious-research/FrescoArchive.research-companion-indexsynthetic-commit-msg-edits
✍️ Commit Message Edits Dataset - 🤖Synthetic
This dataset is a synthetic extension of our expert-labeled commit message edits dataset presented in Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings.
You can check Synthetic tab in our visualization app to browse through the datapoints!
Dataset Structure
Default
Default split contains the synthetic messages generated from expert-labeled dataset by an LLM.… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/synthetic-commit-msg-edits.stego-bench
TODO (maintainer): confirm the final license (the license: field above is a placeholder) and update the LICENSE file before flipping the gate to public. The rest of this card assumes the repo stays gated: manual.
StegoBench
Evaluating steganography potential in language models through supervised learning.
This repository hosts the artifacts that accompany the paper StegoBench: Evaluating steganography potential in language models through supervised learning (NeurIPS 2026… See the full description on the dataset page: https://huggingface.co/datasets/Poseidon-Research/stego-bench.model-toolcall-research
Model Toolcall Research
This dataset stores newline-delimited agent traces from bounded research runs on model repository tool-schema support.
The Dataset Viewer is configured to index only .jsonl files:
toolcall_traces loads trace files under traces/**/*.jsonl.
research_session loads top-level provenance/session traces from *.jsonl.
The archive/ directory preserves the earlier .trace.json uploads for reference, but those files are newline-delimited JSON streams rather than… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/model-toolcall-research.quantum-worldline-research
Quantum Worldline Research Data
Structured research data from the Quantum Worldline project - an AI-assisted research program investigating holographic forces in MERA tensor networks, worldline path integrals on AdS spacetime, and quantum simulation of lattice gauge theories.
Dataset Description
This dataset contains the complete structured output of the Quantum Worldline multi-agent research system, which automates the research cycle: discover - hypothesize - gate - test… See the full description on the dataset page: https://huggingface.co/datasets/Jonboy648/quantum-worldline-research.REval
REval: Reasoning Runtime Behavior of a Program with LLM
Disclaimer: We are not the authors of the REval benchmark. This upload is a convenience repackaging of the original dataset with precomputed execution traces, variable states, and ground truth answers to make the benchmark easier to use programmatically. The original benchmark was created by Junkai Chen et al. and is available at github.com/r-eval/REval. Please cite the original paper if you use this data.
REval is a… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/REval.
