datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open_government
Open Government Dataset
Open Government is the largest agregation of governement text and data made available as part of open data programs.
In total, the dataset contains approximately 380B tokens. While Open Government aims to become a global resource, in its current state it mostly features open datasets from the US, France, European and international organizations.
The dataset comprises 16 collections curated through two different initiaties: Finance commons and Legal commons.… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/open_government.agent-trajectory-sentinel
AgentTrajectorySentinel — 3581 agent episodes across 33 corpora
Committed agent trajectories with step-level telemetry, used to fit and
evaluate one-class monitors for real-time failure detection.
Paper: https://arxiv.org/abs/2608.02464
Code and the full evaluation harness:
https://github.com/sunnydubey1111/agent-trajectory-sentinel
A recorded walkthrough of the method, ending with the live demo
detecting and repairing a real failure:
https://youtu.be/a05n_000klE?t=0… See the full description on the dataset page: https://huggingface.co/datasets/sunnydubey1111/agent-trajectory-sentinel.agent-llm-traces-v2
Exgentic Agent LLM Traces v2 — Agent Chat Only
OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it.
This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.trustworthy-biology-agents-traces
Trustworthy Biology Agents — Run Traces
Raw execution traces from 1,329 agent runs across three coding agents on three
biology benchmarks — BiomniBench-DA, BixBench, and CompBioBench. This is the scrubbed
trace bundle for the study in
manu-tej/ai-scientists; the write-up
lives in that repo's RESULTS.md.
The motivating question is not only whether an agent reaches the right answer, but
whether it behaves like a trustworthy analyst when the task is ambiguous,
under-specified, or… See the full description on the dataset page: https://huggingface.co/datasets/amanutej/trustworthy-biology-agents-traces.agent-traces
Trace Commons — Agent Traces
Trace Commons is one open, public dataset of coding-agent sessions — the
back-and-forth between a developer and an AI coding agent, including prompts,
model responses, tool calls, and command output — contributed voluntarily as an
open resource for studying, evaluating, and building on how these agents
actually work.
Every trace here was donated only from a public, open-source repository, was
anonymized on the contributor's own machine before upload… See the full description on the dataset page: https://huggingface.co/datasets/trace-commons/agent-traces.agentlogs
AgentLogs
AgentLogs is a dataset of activity related to the GitHub agents functionality: repository metadata, agent tasks, sessions, session logs (messages, tool calls, usage details, etc.), and user records.
This dataset is described in:
Jonan Richards, Kosei Horikawa, Youmei Fan, Yutaro Kashiwa, and Mairieli Wessel (2026), AgentLogs: A Dataset for Opening the Black Box of GitHub's Cloud Agent. arXiv: 2608.29204 (preprint).
# Records
Size
Table
Content… See the full description on the dataset page: https://huggingface.co/datasets/risenlab/agentlogs.lmcache-agentic-traces
LMCache Agentic Dataset Collection
A curated dataset collection of 787 multi-turn agentic LLM sessions (24,881 total LLM iterations) designed for benchmarking stateful LLM serving systems. Every session exhibits at least 5 turns with prefix growth and builds to at least 10K tokens of context — making it ideal for evaluating tiered KV Cache solutions like LMCache.
Motivation
Modern LLM agents (coding assistants, research agents, tool-calling systems) make dozens of… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/lmcache-agentic-traces.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.agentic-polymarket
agentic-polymarket
38,915 settled Polymarket binary event markets with full hourly price curves, question text, resolution
terms, and ground truth outcomes. Prepared for research on "getting LLM agents to trade on prediction markets."
Companion code (backtest env + agent trading interface): see RSI-economy/shadow-market.
What this dataset solves
Historical price series cannot be used directly as backtest targets —— a recording does not react to
agent behavior: any… See the full description on the dataset page: https://huggingface.co/datasets/tennant/agentic-polymarket.agent-llm-traces
Multi-Benchmark LLM Agent Traces
A comprehensive dataset of OpenTelemetry traces capturing LLM inference behavior across multiple agent frameworks, benchmarks, and model providers. This dataset enables research into LLM performance analysis, agent behavior patterns, and inference optimization.
Collected by Exgentic - A platform for LLM observability and performance optimization.
Dataset Overview
This dataset contains 1,781 execution traces capturing detailed agent… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces.deepseek-v4-pro-0813-agentic
DeepSeek-V4-Pro 0813 Agentic (DS4)
A standalone, verifiable-first agentic training corpus: 19,072 training traces
plus 2,135 held-out evaluation rows (validation 1,070 / test 1,065), generated by
DeepSeek-V4-Pro 0813 (deepseek-v4-pro-0813, official API, thinking mode) across 13 verifiable task families,
each row admitted only after passing a deterministic programmatic verifier. The corpus is
designed to be directly usable for SFT, GRPO/RLVR, and NeMo Gym / NeMo RL
(verified… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/deepseek-v4-pro-0813-agentic.agent-usage
Agent Usage on the Hugging Face Hub
Coding agents are real users of the Hugging Face Hub. Claude Code, Codex, Cursor, and a growing list of harnesses are searching for models, building and pushing datasets, training models on Jobs, spinning up Spaces — tens of millions of requests so far (hf CLI for agents). Now there's public data on which ones.
Requests made through the huggingface_hub library (including the hf CLI) carry an agent/<name> User-Agent token identifying the… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/agent-usage.agent-collusion
Emergent Collusion in Long-Horizon LLM Agent Interaction
Xinrui Shi*, Yanzhe Zhang*, Diyi Yang
📄 Paper | 💻 Code | 🤗 Data | 🔍 Data Viewer
*Equal contribution.
The experiments reported in the paper and its appendices: 53 conditions, 2,650 trajectories, and 27,100 episodes, of which 600 are warm-up and 26,500 are evaluation episodes. Every condition runs the same 50 fixed task sequences.
Contents
Config / directory
Unit
Count
episodes
One two-agent… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/agent-collusion.agentnet-partial-and-fail-v1
agentnet-partial-and-fail-v1
GUI state transitions (s, a, s') walked on an Ubuntu desktop by
Qwen3.8-27B, from tasks taken from AgentNet and run inside OSWorld's
Docker environment.
Walks the judge ruled partial or failed. These are the larger half and, for a world model, the more useful one: a walk that did not finish still opened dialogs, switched tabs and changed settings, and each of those is a real transition. Measured over eighty-eight walks, a failure visits 13.5 distinct… See the full description on the dataset page: https://huggingface.co/datasets/gui-wm/agentnet-partial-and-fail-v1.Agentic-Long-Context-Understanding-QA 📖 Agentic Long Context Understanding 📖
Self-Taught Agentic Long Context Understanding (Arxiv).
AgenticLU refines complex, long-context queries through self-clarifications and contextual grounding, enabling robust long-document understanding in a single pass.
Installation Requirements
This codebase is largely based on OpenRLHF and Helmet, kudos to them.
The requirements are the same
pip install openrlhf
pip install -r ./HELMET/requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/yzhuang/Agentic-Long-Context-Understanding-QA.ViZDoom-Agentic-Rollouts
ViZDoom Agentic Rollouts
This dataset contains model-generated evaluation results for the pufanyi/ViZDoom benchmark. The benchmark repository contains only test specifications; this repository contains scores and rollout artifacts.
Both configs contain all 120 evaluated episodes: 10 seeds for each of the 12 default single-player ViZDoom environments. Browse the synchronized videos in the ViZDoom Agentic Demo.
Configs
qwen3.6-27b-5tics
The… See the full description on the dataset page: https://huggingface.co/datasets/pufanyi/ViZDoom-Agentic-Rollouts.data-gouv-datasets-catalog
📢 Sondage 2026 : Utilisation des datasets publiques de MediaTech
Vous utilisez ce dataset ou d’autres datasets de notre collection MediaTech ? Votre avis compte !
Aidez-nous à améliorer nos datasets publiques en répondant à ce sondage rapide (5 min) : 👉 https://grist.numerique.gouv.fr/o/albert/forms/gF4hLaq9VvUog6c5aVDuMw/11
Merci pour votre contribution ! 🙌
🇫🇷 Data.gouv.fr Datasets Catalog
This dataset contains a processed and embedded version of the… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/data-gouv-datasets-catalog.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.ninja-agent-traces
Tau retired-king tasks and rollouts
This dataset is written by the Tau validator when a challenger becomes king.
tasks contains one viewer-friendly row per generated task.
rollouts contains one viewer-friendly row per terminal qualification or duel
solve and is the default table shown on the dataset page.
events contains one flattened row per redacted proxy-observed LLM call.
payloads contains complete solution diffs plus request and response bodies split
into bounded, ordered… See the full description on the dataset page: https://huggingface.co/datasets/Wejh/ninja-agent-traces.agentic-coding-trajectories
agentic-coding-trajectories
A unified, tokenized corpus of 15,000 multi-turn agentic-coding sessions (618K turns, 41 turns/session avg) drawn from three publicly-released upstream datasets. Built for benchmarking LLM serving systems on realistic multi-turn coding-agent workloads.
Why this exists
Most LLM serving benchmarks use single-shot prompts. Real coding agents work in long multi-turn loops where each turn appends to a growing prompt. This corpus captures that shape… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/agentic-coding-trajectories.H2EPR-Bench
H²EPR-Bench
An Evidence-Traceable Benchmark for Event-Process Reconstruction
H²EPR-Bench asks a demanding question: can a model reconstruct how a complex
real-world event unfolded, rather than merely summarize what happened? Given an
event specification and fixed multi-source evidence, a system produces a
hierarchical heterogeneous Event-Process Graph (EPG) that makes stages,
episodes, participants, actions, outcomes, relations, and evidence support
explicit.… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/H2EPR-Bench.data-agent
📈 Data Agent
Data-analysis tasks as a plain, load-and-go dataset — no runtime, no framework required. Each
row is one self-contained task: a real tabular dataset, a question about it, and a
deterministically-checkable gold answer. Load it, prompt any model however you like, and grade the
result with the bundled grader.
Where it comes from
Built from the jupyter-agent dataset
— real data-science notebooks over Kaggle datasets. Every question–answer pair was… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/data-agent.Agentic-SLS-Telemetry
Inova-Mk1-Telemetry
Time-aligned printer-state recordings from Inova Mk1 SLS 3D print runs. One row per 10 Hz tick — the recorder's /state/snapshot poll — with the full sensor state snapshot (~64 columns: temperatures, position, power, lights) on every row, the nearest camera frame embedded inline when one fell within the prior 100 ms window, and any 1 kHz position-stream samples from that window collected as a nested list.
25 parquet files across builds spanning 2026-05 through… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/Agentic-SLS-Telemetry.kwaiklear-sample-level-agent-trajectories-2.2Mdole
📢 Sondage 2026 : Utilisation des datasets publiques de MediaTech
Vous utilisez ce dataset ou d’autres datasets de notre collection MediaTech ? Votre avis compte !
Aidez-nous à améliorer nos datasets publiques en répondant à ce sondage rapide (5 min) : 👉 https://grist.numerique.gouv.fr/o/albert/forms/gF4hLaq9VvUog6c5aVDuMw/11
Merci pour votre contribution ! 🙌
🇫🇷 French Legislative Dossiers Dataset (DOLE)
This dataset provides a semantic-ready, chunked and… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/dole.GitHub-Agentic-PR-Dataset
GitHub Agentic PR Dataset
A large-scale dataset of ~2 million GitHub Pull Requests authored by AI coding agents (Claude Code, Cursor, GitHub Copilot, Devin) and human developers — complete with commits, file-level diffs, patches, and bug-fix classification.
The GitHub Agentic PR Dataset is a research-grade corpus for studying how AI coding agents contribute to real-world open-source software, and how their pull requests compare to those written by humans. It pairs 1,959,649 pull… See the full description on the dataset page: https://huggingface.co/datasets/mabujadallah/GitHub-Agentic-PR-Dataset.deepseek-v4-pro-agent-tool-calling-trajectory
DeepSeek V4 Pro ToolScale Agent SFT Dataset
A curated subset of multi-turn tool-calling trajectories generated by DeepSeek V4 Pro on ToolScale. The dataset is filtered by action-match score against ground-truth trajectories and is designed for supervised fine-tuning of agentic models on realistic, multi-step tool use.
Each conversation includes natural-language user requests, tool calls, tool observations, assistant reasoning traces, and grounded final responses across five… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/deepseek-v4-pro-agent-tool-calling-trajectory.foldtowel_mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_roll.pos",
"left_gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/CSI-Agent/foldtowel_merged.Qwen-3.6-plus-agent-tool-calling-trajectory
Qwen 3.6 Plus: ToolScale Agent SFT Dataset
Multi-turn tool-calling trajectories generated by Qwen 3.6 Plus via OpenRouter on ToolScale. Both passing and near-passing rollouts are included, allowing users to choose their own quality threshold using reward and score.
Each row is a flattened conversation prefix ending at one assistant turn, ready for next-token SFT. Assistant turns include a reasoning_content field containing the model’s reasoning.
What's inside… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen-3.6-plus-agent-tool-calling-trajectory.Agentic_Movielens
Dataset Card for Agentic_Movielens
Dataset Description
This dataset contains movie ratings and related information.
Usage
Load the dataset using the HuggingFace datasets library:
from datasets import load_dataset
dataset = load_dataset("ShuzeChen/Agentic_Movielens")
Dataset Structure
The dataset is provided in the train split and includes all collected data.
Additional Information
For questions or issues, please refer to the repository… See the full description on the dataset page: https://huggingface.co/datasets/ShuzeChen/Agentic_Movielens.
