datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.AgenticDataBench
AgenticDataBench: A Comprehensive Benchmark for Data Agents
Project Page | GitHub | Paper
AgenticDataBench is a comprehensive benchmark for evaluating LLM-based data agents that automate real-world data science workflows. It addresses the lack of rigorous evaluation by providing diverse, realistic tasks with fine-grained ground-truth labels.
The benchmark spans 15 domains, including real B2B fintech use cases, and is structured around reusable data science skills—core… See the full description on the dataset page: https://huggingface.co/datasets/shawnzzzh/AgenticDataBench.sol-max-opusnode-data
sol-max-opusnode-data
Training data built by the AgentPTB arm for cell sol-max-opusnode — Codex / gpt-5.6-sol @ effort max.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/sol-max-opusnode.h*, and the companion to the run record in agentic-ptb/sol-max-opusnode-record.
field
value
plot cell
sol-max-opusnode
driver
Codex / gpt-5.6-sol… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/sol-max-opusnode-data.sol-max-data
sol-max-data
Training data built by the AgentPTB arm for cell sol-max — Codex / gpt-5.6-sol @ effort max.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/sol-max.h*, and the companion to the run record in agentic-ptb/sol-max-record.
field
value
plot cell
sol-max
driver
Codex / gpt-5.6-sol
reasoning effort
max
total size
114.47 GB… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/sol-max-data.EMU-Agentic-PostTrain-Dataopus-high-v3-data
opus-high-v3 — complete research record
This dataset archives the qualitative and quantitative record of the
msr-agentic-ptb-opus / opus-high-v3 Claude Code research run.
The submitted artifact uses the unmodified base weights with a two-attempt
Pi verifier harness. The final replicated SWE result was 24.6% (245/995) with
the stock scaffold and 32.4% (321/990) with the submitted harness. Training
did not improve the weights; all trained variants measured at or below the
base… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/opus-high-v3-data.sol-max-v2-dataagentic-datasetGitHub-Agentic-PR-Dataset
GitHub Agentic PR Dataset
A large-scale dataset of ~2 million GitHub Pull Requests authored by AI coding agents (Claude Code, Cursor, GitHub Copilot, Devin) and human developers — complete with commits, file-level diffs, patches, and bug-fix classification.
The GitHub Agentic PR Dataset is a research-grade corpus for studying how AI coding agents contribute to real-world open-source software, and how their pull requests compare to those written by humans. It pairs 1,959,649 pull… See the full description on the dataset page: https://huggingface.co/datasets/mabujadallah/GitHub-Agentic-PR-Dataset.agentic-tool-call-dataset-12k
Agentic Tool Calling Dataset 12K
A curated 12K-sample tool-calling SFT dataset in a TRL-ready chat format. Each sample contains multi-turn agent trajectories with explicit reasoning, structured tool_calls, and tool responses.
Dataset Summary
Property
Value
Total Samples
12,000
Short split
10,000 (agent_short_10k.jsonl)
Long split
2,000 (agent_long_2k.jsonl)
Language
English
Format
OpenAI-style messages with tool_calls
License
Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/pyromind/agentic-tool-call-dataset-12k.agentic-midtrain-data
Agentic Midtraining Trajectories
Four exact-deduplicated training artifacts in the shared
agentic_trajectory_v1 schema. There are two alternative category mixtures,
each available with compacted or complete redacted tool outputs. Each Parquet
row is one complete trajectory.
Config
Included categories
Excluded category
Rows
Nemotron post-template tokens
no_cyber
SWE, General Coding, Terminal Use
Cyber
846,367
19,572,336,927
no_general_coding
SWE, Terminal Use, Cyber… See the full description on the dataset page: https://huggingface.co/datasets/secmlr/agentic-midtrain-data.dpsk-v4-flash-data
dpsk-v4-flash-data
Training data built by the AgentPTB arm for cell dpsk-v4-flash — pi / DeepSeek v4-flash @ effort thinking.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/dpsk-v4-flash.h*, and the companion to the run record in agentic-ptb/dpsk-v4-flash-record.
field
value
plot cell
dpsk-v4-flash
driver
pi / DeepSeek v4-flash… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/dpsk-v4-flash-data.Agentic-Chain-of-Thought-Coding-SFT-Dataset
🤖 Agentic Coding CoT Dataset
A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities.
📋 Dataset Description
This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns.
🏗️ Assistant Data Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Agentic-Chain-of-Thought-Coding-SFT-Dataset.AgenticDataBench
Paper
This dataset is described in:
https://arxiv.org/abs/2607.01647
Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yuqing1207/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.agentic-mqa-objectsopus-max-data
opus-max-data
Training data built by the AgentPTB arm for cell opus-max — Claude Code / claude-opus-5 @ effort max.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/opus-max.h*, and the companion to the run record in agentic-ptb/opus-max-record.
field
value
plot cell
opus-max
driver
Claude Code / claude-opus-5
reasoning effort
max… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/opus-max-data.glm52-datagen-r11-100-agentic-function-calling-pivot-v2-tracesAgenticData_5kgrug-67b-a2b-agentic-sft-training-data
Grug 67B agentic SFT training dataset
This directory is a local, revision-pinned reconstruction of the exact 29-component mixture consumed by grug_67b_a2b_sft_s3_agentic.
The reconstruction has two representations:
converted_hf/ contains the readable converted datasets. Each component is checked out at the full Hugging Face commit recorded by the corresponding Marin document artifact. These 29 snapshots contain 77,012 conversations and occupy about 1.67 GB before filesystem… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/grug-67b-a2b-agentic-sft-training-data.agentic-publication-protocol-dataset
APP compare-app benchmark
Paired reader conversations and blinded evaluations comparing an Agentic
Publication Protocol (APP) paper agent against a general repository-aware
agent, on 11 quantum-physics papers.
For each paper, a neutral reader asks the same scripted questions to both agents;
the two transcripts are anonymized and scored by a blinded evaluator on
accuracy, informativeness, grounding, and honesty (1-10).
Evaluator: Codex CLI, gpt-5.5, reasoning effort xhigh… See the full description on the dataset page: https://huggingface.co/datasets/phynics/agentic-publication-protocol-dataset.agentic_code_dataset_22Dataset: 22 Real Claude Code Sessions
To validate Suffix Decoding's applicability in Agentic Coding scenarios, we collected 22 complete Claude Code session recordings.
Dataset Overview
Metric
Value
Collection date
December 2025
Total sessions
22
Total conversation turns
17,487
Total runtime
50 hours
Total input tokens
6,996,619
Total output tokens
6,094,906
Session Scale Distribution
Statistic
Min
Max
Average
Conversation turns
273… See the full description on the dataset page: https://huggingface.co/datasets/novita/agentic_code_dataset_22.agentic-data-pipeline-10kglm52-datagen-r11-101-agentic-indirect-prompt-injection-v2-tracesAgentic-SLS-Database
Agentic-SLS-Database
Canonical graph dataset of Inova Mk1 SLS printer entities: jobs, print sessions, print profiles, and objects (STL geometry). Each entity is its own HF config; relationships are encoded as ID references between rows.
Domain-specific datasets (e.g. ppak10/Agentic-SLS-ASTM) reference rows here by ID and may embed frozen snapshots of the referenced state.
Configs
Config
Description
Script
Output
jobs
One row per .s4a print job, with… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/Agentic-SLS-Database.agentic-critic-dataset
Agentic Critic Dataset
High-quality AIGC images with rich metadata for aesthetic evaluation.
Metadata Fields
Each entry in metadata.jsonl contains:
prompt: Positive prompt
negative_prompt: Negative prompt
model: Model name and hash
sampler: Sampling method
steps: Generation steps
cfg_scale: CFG scale
seed: Random seed
stats: Engagement metrics
image_path: Relative path to image
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/ChengyouJia/agentic-critic-dataset.grok-data
grok-data
Training data built by the AgentPTB arm for cell grok — pi / grok-4.6 @ effort xhigh.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/grok.h*, and the companion to the run record in agentic-ptb/grok-record.
field
value
plot cell
grok
driver
pi / grok-4.6
reasoning effort
xhigh
total size
2.54 GB
path in run… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/grok-data.Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1
🤖 Agentic Coding CoT Dataset v1.1
A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities.
📋 Dataset Description
This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 & MiniMax M2.1 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns.
🏗️ Assistant… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1.sol-high-data
sol-high-data
Training data built by the AgentPTB arm for cell sol-high — Codex / gpt-5.6-sol @ effort high.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/sol-high.h*, and the companion to the run record in agentic-ptb/sol-high-record.
field
value
plot cell
sol-high
driver
Codex / gpt-5.6-sol
reasoning effort
high
total size… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/sol-high-data.github-agentic-workflows-dataset
GitHub Agentic Workflows Dataset
Este dataset contiene la información extraída y estructurada de archivos Markdown (.github/workflows/*.md) pertenecientes a repositorios de GitHub que implementan Agentic Workflows.
Estructura del Dataset
El dataset se organiza en 3 tablas relacionales en formato Apache Parquet:
repositories.parquet: Identificadores y nombres de los repositorios procesados.
workflow_files.parquet: Archivos .md localizados, con su ruta original y… See the full description on the dataset page: https://huggingface.co/datasets/ry-02/github-agentic-workflows-dataset.
