datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.agentic-ai-options-resultsagentic-redteam-benchmark
agentic-redteam-benchmark
v0.8 preview · 2,288 multi-step agent trajectories · 513 hand-authored gold + 1,775 provenance-flagged augmented.
A per-step benchmark that scores whether a verifier catches drift inside an agent's trajectory — not whether a prompt is harmful.
📦 Code, eval harness & issues: github.com/Alkur123/agentic-redteam-benchmark · 📄 Paper: A Per-Step Trajectory Benchmark for AI-Agent Governance Verifiers and a Corrected Catch-at-Drift Metric (Aegis AI, 2026)… See the full description on the dataset page: https://huggingface.co/datasets/jash-ai/agentic-redteam-benchmark.TR-HASH-Pretraining-125B-Agentic-32K
TR-HASH Pretraining 125B — Agentic 32K
Private, source-curated pretraining artifact for the TR-HASH Agentic 32K model
line. It contains 125B packed token exposures: 75B foundation and 50B
agentic/procedural content.
The corpus uses the immutable, validated 32,000-ID revision of
AETHORIA-AI/TR-HASH-Tokenizer-32K-Agentic. It is not compatible with the
older TR-HASH 32K tokenizer.
Composition
Bucket
Tokens
Purpose
Foundation
75B
English and French… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/TR-HASH-Pretraining-125B-Agentic-32K.Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yuqing1207/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.daily-oracle
Daily Oracle
📰 Project Website📝 Paper - Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
Daily Oracle is a continuous evaluation benchmark using automatically generated QA pairs from daily news to assess how the future prediction capabilities of LLMs evolve over time.
Dataset Details
Question Type: True/False (TF) & Multiple Choice (MC)
Current Version*
Time Span: 2020.01.01 - 2026.07.18
Size: 20,376 TF questions and 18,557 MC… See the full description on the dataset page: https://huggingface.co/datasets/agentic-learning-ai-lab/daily-oracle.TR-HASH-Agentic-SFT-32K-210K
TR-HASH Agentic SFT 32K
Balanced instruction and tool-use SFT data for
AETHORIA-AI/TR-HASH-Tokenizer-32K-Agentic.
The canonical repository name is retained, while its contents replace the former
tool-heavy 21K laboratory corpus.
Composition
Split
General instruction
Tool-aware
Total
Train
182,000
18,000
200,000
Validation
9,000
1,000
10,000
The 9% tool-aware training slice contains tool calls, no-call decisions with
tools present, and final… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/TR-HASH-Agentic-SFT-32K-210K.daily-paper-2026-09-11-reasoning-verbosity-tax-agentic-cost
The Reasoning Verbosity Tax: Per-Turn Cost-Quality Frontiers of Explicit vs. Compact Chain-of-Thought in Self-Hosted Agentic LLMs on H200
TL;DR — An analytical paper that turns the Qwen3 thinking toggle into a priced dial for self-hosted agentic LLM loops: explicit chain-of-thought persisted in context compounds a per-turn verbosity tax quadratically in the horizon (Theta(T^2)) versus linearly (Theta(T)) under a discard policy, the relative tax widens monotonically toward a… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-09-11-reasoning-verbosity-tax-agentic-cost.Got_Agentic_AI_5k
Got_Agentic_AI_5k
A 5,000-example dataset to train LLMs into production-grade agentic assistants (“Angelic Agents”): high-agency, tool-aware, test-driven, and safety-first.
This dataset focuses on the kinds of tasks real engineering teams and major AI developers care about:
Diff-first coding patches and tests
Planner–executor agent architectures
Evals, monitoring, and rollback discipline
Data engineering transforms with quality checks
Incident postmortems and operational… See the full description on the dataset page: https://huggingface.co/datasets/11-47/Got_Agentic_AI_5k.OWASP-Agentic-AI-Threats
OWASP Agentic AI Threats and Mitigations
Dataset Summary
The OWASP Agentic AI Threats and Mitigations Question Answering Dataset is a
synthetic instruction-style question-answering dataset derived from the OWASP
Agentic AI - Threats and Mitigations report.
The dataset is designed to support training, fine-tuning, retrieval evaluation,
and domain-specific question-answering use cases related to agentic AI security,
large language model agents, multi-agent… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/OWASP-Agentic-AI-Threats.TR-HASH-Agentic-SFT-2M
TR-HASH Agentic SFT 2M
Clean full-parameter SFT corpus for the TR-HASH 100M Agentic line. This build contains 2,000,000 training examples and 100,000 validation examples. It does not replay the previous 250K or 500K SFT releases.
Composition
Component
Train
Validation
general_conversations
800,000
40,000
verified_compact_reasoning
450,000
22,500
python_code
350,000
17,500
constraints_and_formats
150,000
7,500
tool_available_direct_answer
150,000… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/TR-HASH-Agentic-SFT-2M.Execution-Finality-Security-for-Agentic-AI-Autonomous-Systems-Cloud-Payments-Telecom-OS-and-Robo
Dataset Description
The architecture addresses a structural gap in modern AI and autonomous systems: the separation between computation and external consequence. Existing protocols and controls (identity, access control, encryption, logging, policy engines) govern movement, authentication, and recording of data. They do not, by themselves, make the transition from a generated act to an externally effective act a protected technical precondition.
This dataset provides a clean… See the full description on the dataset page: https://huggingface.co/datasets/sangamdas/Execution-Finality-Security-for-Agentic-AI-Autonomous-Systems-Cloud-Payments-Telecom-OS-and-Robo.Walking-Tours-Semantic
Walking Tours Semantic
Walking Tours Semantic (WT-Sem), introduced in PooDLe, provides semantic segmentation masks for videos in the Walking Tours dataset, as well as three additional videos for validation.
Frames are sampled every 2 seconds from each video and a top-of-the-line semantic segmentation model, OpenSeed, is used to generate the masks.
Specifically, the Swin-L variant of OpenSeed, pretrained on COCO and Objects365 and finetuned on ADE20K, is used.
The 3 new walkaround… See the full description on the dataset page: https://huggingface.co/datasets/agentic-learning-ai-lab/Walking-Tours-Semantic.ai-research-berkeley-agentic-verification-harness-opt
Agentic Verification Meta-Verifier Traces
This public, manually gated Dataset repository stores immutable phase snapshots
from Meta-Verifier experiments.
Each run is organized as:
experiments/<theme>/<method>/<run>/phases/
train/ # Solver, delegated-verifier, Proposer/Reflector, harness population
val/ # Full validation traces, metrics, and selected frozen harness
heldout/ # Claimed held-out stage, all scheduled cells, and final scores
Access requests are… See the full description on the dataset page: https://huggingface.co/datasets/MinjaeLee-FuriosaAI-Ext/ai-research-berkeley-agentic-verification-harness-opt.Agentic-aihendar-agentic-ai-dataset
Hendar Agentic AI Evaluation & Security Benchmark
A compact, expert-authored benchmark for evaluating trustworthy agentic AI systems across capability, tool use, retrieval, security, policy enforcement, multi-agent coordination and regression safety.
This dataset is a public companion to the Agentic AI Academy by Hendar Mawan, PhD. It is designed for evaluation, CI regression testing, red-team exercises and engineering education—not as a generic instruction-tuning corpus.… See the full description on the dataset page: https://huggingface.co/datasets/h0000w/hendar-agentic-ai-dataset.TR-HASH-Agentic-SFT-32K-500K
TR-HASH Agentic SFT 32K 500K
Audited 500K-example long-horizon SFT mixture for the native TR-HASH Agentic
tokenizer. It cleans the pinned 250K corpus and adds verified examples for
verified arithmetic, compact mathematical reasoning, calculator use,
execution-checked code, constrained instruction following, and general answers.
Training content
Supervised behavior
Examples
Share
General dialogue and instruction responses
Answer diverse questions and follow ordinary… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/TR-HASH-Agentic-SFT-32K-500K.APP1-Agentic-Safety-SFT-Datadaily-paper-2026-09-20-agentic-prefix-reuse-manifest-ordering
The Manifest Order: Measuring KV-Cache Prefix Reuse and Token Cost Savings from Skill-Manifest and Tool-Schema Ordering in Unattended Agentic Loops on H200
TL;DR — A closed-form model of the prefix-cache economics of the fixed prompt prefix (system prompt + tool schemas + skill manifest) that unattended agent loops re-send every turn. Manifest ordering is a zero-training, deployment-time lever that, under the model's stated assumptions, cuts input-token cost by 12-89% under… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-09-20-agentic-prefix-reuse-manifest-ordering.MOSAIC-agentic-3m
Agent Activity Dataset
This dataset is released in conjunction with the paper Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time, accepted at MSR 2026.
Dataset Overview
The dataset contains a total of 111,969 Pull Requests (June through August 2025) from both coding agents (Claude Code, OpenAI Codex, GitHub Copilot, Google Jules, and Devin) and human contributors. It also includes additional activity metadata such as… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/MOSAIC-agentic-3m.GSM8K-TR-HASH-Agentic-32K
GSM8K TR-HASH Agentic 32K
Public experimental projection of the pinned GSM8K train split for TR-HASH Agentic SFT. The GSM8K test split is excluded from every training file and reserved for supervised-checkpoint evaluation.
Source: openai/gsm8k @ 740312add88f781978c0658806c59bc2815b9866
Tokenizer: AETHORIA-AI/TR-HASH-Tokenizer-32K-Agentic @ 2fcbc2c5359ded0244ca14531f1b3806eebac55e
Train examples: 7,473
Training epochs are a model-run choice; this artifact contains one copy of… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/GSM8K-TR-HASH-Agentic-32K.Alexander-Agentic
A dataset for creating agentic models
Dataset Summary
Alexander-Agentic contains agentic traces generated by frontier models, extracted using AI harnesses such as Pi, Codex, and Claude Code. Each example is formatted following the Transformers messages schema, ready for fine-tuning agentic models.
Domains covered: coding, research, tool use, etc.
Source harnesses: Pi, Codex, Claude Code
Format: OpenAI/Transformers messages schema
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Aquiles-ai/Alexander-Agentic.agentic-web-cheatsheets
Bowmark: Agentic Web Cheatsheets — Free Sample
Bowmark indexes how websites actually work, for AI agents.
Each row is a cheatsheet for one task on one site: the behavioral gotchas you
only learn by driving the site, a deep-link shortcut where one exists, and a
verification stamp saying how many times it worked and as of when. Every row
was run end-to-end and proven to work — that's the gate to be included.
This repository is a free, curated sample — the strongest… See the full description on the dataset page: https://huggingface.co/datasets/bowmark-ai/agentic-web-cheatsheets.Portfolio-OptimizationFinancial-Advisory-ClientsNemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The… See the full description on the dataset page: https://huggingface.co/datasets/meet-the-1337/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/supraja04/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.sleeetview_agentic_ai_dataset
The SleetView Agentic AI Dataset
The SleetView Agentic AI dataset is a collection of synthetic content automatically generated using Agentic AI
Dataset Details
Dataset Description
The images were generated with a collection of models available under the Apache-2.0 or creativeml-openrail-m licenses.
To generate this dataset we used our own agentic implementation given the goal of creating a dataset that can be used to research synthetic content detection.
As… See the full description on the dataset page: https://huggingface.co/datasets/ninamoss/sleeetview_agentic_ai_dataset.ai-research-berkeley-agentic-verification-benchmarkInternal documentation
entity-classification-agentic-ai
Entity Classification for Agentic AI Systems
This dataset contains 12,497 samples for entity classification in agentic AI systems.
Splits
train.json: 8,747 samples
validation.json: 1,250 samples
test.json: 2,500 samples
Entity Types
Agent, Task, Tool, Input, Output, Human
Usage
from datasets import load_dataset
dataset = load_dataset("holistic-ai/entity-classification-agentic-ai")
Fields
content: Text to classify
expected_entity:… See the full description on the dataset page: https://huggingface.co/datasets/holistic-ai/entity-classification-agentic-ai.
