datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.agentic-ai-options-resultsagentic-redteam-benchmark
agentic-redteam-benchmark
v0.8 preview · 2,288 multi-step agent trajectories · 513 hand-authored gold + 1,775 provenance-flagged augmented.
A per-step benchmark that scores whether a verifier catches drift inside an agent's trajectory — not whether a prompt is harmful.
📦 Code, eval harness & issues: github.com/Alkur123/agentic-redteam-benchmark · 📄 Paper: A Per-Step Trajectory Benchmark for AI-Agent Governance Verifiers and a Corrected Catch-at-Drift Metric (Aegis AI, 2026)… See the full description on the dataset page: https://huggingface.co/datasets/jash-ai/agentic-redteam-benchmark.Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yuqing1207/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.daily-oracle
Daily Oracle
📰 Project Website📝 Paper - Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
Daily Oracle is a continuous evaluation benchmark using automatically generated QA pairs from daily news to assess how the future prediction capabilities of LLMs evolve over time.
Dataset Details
Question Type: True/False (TF) & Multiple Choice (MC)
Current Version*
Time Span: 2020.01.01 - 2026.07.18
Size: 20,376 TF questions and 18,557 MC… See the full description on the dataset page: https://huggingface.co/datasets/agentic-learning-ai-lab/daily-oracle.TR-HASH-Agentic-SFT-32K-210K
TR-HASH Agentic SFT 32K
Balanced instruction and tool-use SFT data for
AETHORIA-AI/TR-HASH-Tokenizer-32K-Agentic.
The canonical repository name is retained, while its contents replace the former
tool-heavy 21K laboratory corpus.
Composition
Split
General instruction
Tool-aware
Total
Train
182,000
18,000
200,000
Validation
9,000
1,000
10,000
The 9% tool-aware training slice contains tool calls, no-call decisions with
tools present, and final… See the full description on the dataset page: https://huggingface.co/datasets/AETHORIA-AI/TR-HASH-Agentic-SFT-32K-210K.Got_Agentic_AI_5k
Got_Agentic_AI_5k
A 5,000-example dataset to train LLMs into production-grade agentic assistants (“Angelic Agents”): high-agency, tool-aware, test-driven, and safety-first.
This dataset focuses on the kinds of tasks real engineering teams and major AI developers care about:
Diff-first coding patches and tests
Planner–executor agent architectures
Evals, monitoring, and rollback discipline
Data engineering transforms with quality checks
Incident postmortems and operational… See the full description on the dataset page: https://huggingface.co/datasets/11-47/Got_Agentic_AI_5k.hendar-agentic-ai-dataset
Hendar Agentic AI Evaluation & Security Benchmark
A compact, expert-authored benchmark for evaluating trustworthy agentic AI systems across capability, tool use, retrieval, security, policy enforcement, multi-agent coordination and regression safety.
This dataset is a public companion to the Agentic AI Academy by Hendar Mawan, PhD. It is designed for evaluation, CI regression testing, red-team exercises and engineering education—not as a generic instruction-tuning corpus.… See the full description on the dataset page: https://huggingface.co/datasets/h0000w/hendar-agentic-ai-dataset.OWASP-Agentic-AI-Threats
OWASP Agentic AI Threats and Mitigations
Dataset Summary
The OWASP Agentic AI Threats and Mitigations Question Answering Dataset is a
synthetic instruction-style question-answering dataset derived from the OWASP
Agentic AI - Threats and Mitigations report.
The dataset is designed to support training, fine-tuning, retrieval evaluation,
and domain-specific question-answering use cases related to agentic AI security,
large language model agents, multi-agent… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/OWASP-Agentic-AI-Threats.Alexander-Agentic
A dataset for creating agentic models
Dataset Summary
Alexander-Agentic contains agentic traces generated by frontier models, extracted using AI harnesses such as Pi, Codex, and Claude Code. Each example is formatted following the Transformers messages schema, ready for fine-tuning agentic models.
Domains covered: coding, research, tool use, etc.
Source harnesses: Pi, Codex, Claude Code
Format: OpenAI/Transformers messages schema
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Aquiles-ai/Alexander-Agentic.APP1-Agentic-Safety-SFT-DataExecution-Finality-Security-for-Agentic-AI-Autonomous-Systems-Cloud-Payments-Telecom-OS-and-Robo
Dataset Description
The architecture addresses a structural gap in modern AI and autonomous systems: the separation between computation and external consequence. Existing protocols and controls (identity, access control, encryption, logging, policy engines) govern movement, authentication, and recording of data. They do not, by themselves, make the transition from a generated act to an externally effective act a protected technical precondition.
This dataset provides a clean… See the full description on the dataset page: https://huggingface.co/datasets/sangamdas/Execution-Finality-Security-for-Agentic-AI-Autonomous-Systems-Cloud-Payments-Telecom-OS-and-Robo.MOSAIC-agentic-3m
Agent Activity Dataset
This dataset is released in conjunction with the paper Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time, accepted at MSR 2026.
Dataset Overview
The dataset contains a total of 111,969 Pull Requests (June through August 2025) from both coding agents (Claude Code, OpenAI Codex, GitHub Copilot, Google Jules, and Devin) and human contributors. It also includes additional activity metadata such as… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/MOSAIC-agentic-3m.agentic-web-cheatsheets
Bowmark: Agentic Web Cheatsheets — Free Sample
Bowmark indexes how websites actually work, for AI agents.
Each row is a cheatsheet for one task on one site: the behavioral gotchas you
only learn by driving the site, a deep-link shortcut where one exists, and a
verification stamp saying how many times it worked and as of when. Every row
was run end-to-end and proven to work — that's the gate to be included.
This repository is a free, curated sample — the strongest… See the full description on the dataset page: https://huggingface.co/datasets/bowmark-ai/agentic-web-cheatsheets.Portfolio-OptimizationFinancial-Advisory-ClientsNemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/supraja04/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The… See the full description on the dataset page: https://huggingface.co/datasets/meet-the-1337/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.entity-classification-agentic-ai
Entity Classification for Agentic AI Systems
This dataset contains 12,497 samples for entity classification in agentic AI systems.
Splits
train.json: 8,747 samples
validation.json: 1,250 samples
test.json: 2,500 samples
Entity Types
Agent, Task, Tool, Input, Output, Human
Usage
from datasets import load_dataset
dataset = load_dataset("holistic-ai/entity-classification-agentic-ai")
Fields
content: Text to classify
expected_entity:… See the full description on the dataset page: https://huggingface.co/datasets/holistic-ai/entity-classification-agentic-ai.Portfolio-RebalancePersonal-Finance-Dataclaude-semtools-resultspararev-casimir-compound-fixed
ParaRev Full-Context Paragraph Revision Dataset
Overview
This dataset provides paragraph-level revision pairs with access to the full scientific article context.
Each sample includes a full article, a selected paragraph, its revised version, and an instruction prompt, supporting long-context academic revision tasks.
Data Sources
This dataset is derived from publicly available academic datasets:
CASIMIR: multi-version scientific articles used to reconstruct… See the full description on the dataset page: https://huggingface.co/datasets/Edinburgh-AgenticAI/pararev-casimir-compound-fixed.Portfolio-ManagementFinancial-ReportsCredit-Portfolio-OptimizationIntraday-Tradingqwen-agentic-json-thinkingagentic-ai-instructions-id-en-cleaned
🧹 Agentic AI Instructions (ID-EN) - Cleaned Version
This dataset is the 100% cleaned and validated version of Kimsang766/agentic-ai-instructions-id-en. It has been processed through an enterprise-grade automated data pipeline to ensure pristine data quality for AI and LLM instruction-tuning.
🛠️ Cleaning Pipeline (The Process)
This dataset was sanitized using Polars and validated to guarantee:
0 Missing Values (Nulls): All rows with empty translations were… See the full description on the dataset page: https://huggingface.co/datasets/Kimsang766/agentic-ai-instructions-id-en-cleaned.agentic-ai-instructions-id-en
🤖 Agentic AI Instructions (ID-EN) - Synthetic Data Pipeline
📖 What is this Project? (Deskripsi)
This project is an Automated Synthetic Data Generation Pipeline designed to create high-quality, bilingual (English & Indonesian) datasets for training Agentic AI.
Instead of writing data manually, this project uses a local Large Language Model (LLM) running on Ollama to autonomously generate, parse, validate, and compile 1,200 complex instruction-response pairs… See the full description on the dataset page: https://huggingface.co/datasets/Kimsang766/agentic-ai-instructions-id-en.
