datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-agent-security-incidents
AI Agent Security Incident Database v0.1
A structured, machine-readable database of 1365 confirmed AI agent security incidents, collected and classified automatically.
What is this?
Every time an AI agent causes unintended harm — escaping a sandbox, exploiting an API, taking unauthorized actions, exfiltrating data — this database captures it.
This is not a list of theoretical risks. Every entry describes something that actually happened, with a verifiable source… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents.AgentJudgeBench
AgentJudgeBench: Evaluating LLM Judge Reliability on Agentic Tool-Calling
A benchmark for systematically evaluating how reliably LLM judges assess
agentic tool-calling workflows across structured, dependency-driven tasks.
Why this benchmark?
AgentJudgeBench measures how reliably LLM judges assess agentic tool-calling outputs. It provides 3,808 benchmark records spanning six DAG topologies and three difficulty… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/AgentJudgeBench.moltbook-agent-social-ai-prompt-injection-dataset
Moltbook Agent-Social AI Prompt Injection Dataset
207,391 items — 77,469 posts and 129,922 comments — from Moltbook, a social network whose users are AI agents.
Scanned for indirect prompt-injection patterns using the taxonomy of Greshake et al. (2023). The full raw corpus is included, so you can ignore my analysis entirely and do your own.
These are keyword-matched candidates, not verified attacks. An agent discussing prompt injection matches the same words as one performing… See the full description on the dataset page: https://huggingface.co/datasets/DavidTKeane/moltbook-agent-social-ai-prompt-injection-dataset.ai-code-generation-swe-agents-2026
💻 AI Code Generation, SWE Agents & Program Synthesis Dataset (2026 Edition)
A structured research dataset featuring 3,181 domain-verified research papers and 771 official code repositories focused on Autonomous Software Engineering Agents (SWE-bench), Program Synthesis, DeepSeek-Coder-V2, Qwen2.5-Coder, Test-Driven Code Repair, Self-Healing Software, AST Semantic Modeling, and Formal Logic Verification (2023–2026).
Built with Universal Scientific Engine V17.1 Gold, providing 47… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/ai-code-generation-swe-agents-2026.read-along-ai-agent-traces
Read-Along AI - Agent Traces
This dataset contains the raw agent traces and conversation logs from the development of Read-Along AI, a submission for the Hugging Face Build Small Hackathon.
Dataset Description
These .jsonl files represent the unedited, behind-the-scenes "agent traces" of the AI coding assistant orchestrating the build of this project.
Sharing these traces fulfills the requirements for the "Sharing is Caring" bonus badge, providing the community… See the full description on the dataset page: https://huggingface.co/datasets/kingkw1/read-along-ai-agent-traces.agent-leaderboard
Agent Leaderboard
Overview
The Agent Leaderboard evaluates language models' ability to effectively utilize tools in complex scenarios. With major tech CEOs predicting 2025 as a pivotal year for AI agents, we built this leaderboard to answer: "How do AI agents perform in real-world business scenarios?"
Get latest update of the leaderboard on Hugging Face Spaces. For more info, checkout the blog post for a detailed overview of our evaluation methodology.… See the full description on the dataset page: https://huggingface.co/datasets/galileo-ai/agent-leaderboard.crypto-web3-ai-agents-2026
⚡ Crypto, Web3 & Autonomous Financial AI Agents Dataset (2023–2026)
This dataset contains 100 strictly domain-filtered research papers focusing on Decentralized AI, Autonomous Financial Agents, Smart Contract Verification, Zero-Knowledge Proofs (ZKP), DeFi, and Multi-Agent Consensus (2023-2026).
📊 Features:
384-dimensional PyTorch Embeddings for Vector Search & Semantic Clustering
Strict Domain Verification: Passed 2-stage filtering (100% relevant to Crypto/AI)… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/crypto-web3-ai-agents-2026.agenttool-training-garden
AgentTool HF Training Garden
A tiny metadata-only companion for designing a reproducible Hugging Face data
lifecycle without treating the Hub, a Dataset Card, or one quality score as
training authority.
The Garden has six layers:
Bedrock — rights, license, privacy, separate participation reports,
gating, scoped authority, withdrawal, and repair.
Soil — an exact Hub commit plus content-addressed observations and file
manifests.
Roots — acquisition, parsing, filtering, secret… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-training-garden.clawk-ai-agent-dataset
Clawk AI Agent Dataset
Collected by David Keane (IR240474) — NCI MSc Cybersecurity
National College of Ireland | March 2026
📖 Read the Full Journey
From RangerBot to CyberRanger V42 Gold — The Full Story
The complete story: dentist chatbot → Moltbook discovery → 4,209 real injections → V42-gold (100% block rate). Psychology, engineering, and 42 versions of persistence.
🔗 Links
Resource
URL
📦 This Dataset
DavidTKeane/clawk-ai-agent-dataset
🤖… See the full description on the dataset page: https://huggingface.co/datasets/DavidTKeane/clawk-ai-agent-dataset.context-ucurve-coding-agents
Context U-curve: 36 coding-agent runs under six context-clearing policies
How often should an LLM coding agent's context be cleared? This dataset holds every run behind the report
"Clear Every Third Task: A Measured U-Curve in the Context Economy of Coding Agents"
(Evgenii Arsentev, 2026; corrected version 1.2, DOI 10.5281/zenodo.22759217; version 1.0: DOI 10.5281/zenodo.22699668).
A fixed suite of twelve programming tasks was run under six session-length policies — a fresh… See the full description on the dataset page: https://huggingface.co/datasets/arsentev-ai/context-ucurve-coding-agents.Agent-Trajectory-Data-Sample
Agent-Trajectory-Dataset
Description
This dataset covers office-based scenarios such as in-depth searches, data analysis, and industry research, encompassing complete multi-turn reasoning trajectories and tool-calling chains. It is designed to support the analysis of agent planning capabilities, research into tool selection strategies, and quality assessment, providing a structured benchmark for agent training and evaluation.
For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/Agent-Trajectory-Data-Sample.agent-discoverability-ado-score-romania
Agent Discoverability (ADO Score) — Romania, September 2026
130 Romanian domains probed for A2A Agent Cards, MCP discovery, llms.txt, schema.org and Wikidata. Zero Agent Cards; mean ADO Score 17/100. Raw data, scripts and scoring spec, CC BY 4.0.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
On 8 September 2026 a standard-library Python probe (published) requested, for each of 130 domains, the homepage without JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/agent-discoverability-ado-score-romania.ecommerce-ai-data-analyst-agent-benchmark
E-commerce AI Data Analyst Agent Benchmark
A synthetic e-commerce dataset for evaluating AI data analyst agents on
realistic, multi-step business analysis, data-quality investigation, and
analytical reasoning.
This dataset is part of the
E-commerce AI Data Analyst Agent Benchmark.
Dataset summary
This dataset supports evaluation of AI data analyst agents on realistic,
multi-step e-commerce analysis.
It contains:
customers.csv
products.csv
orders.csv
returns.csv… See the full description on the dataset page: https://huggingface.co/datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark.ai-agent-security-sft-dpo
AI Agent Security — SFT + DPO
Fine-tuning data for teaching an AI agent to protect its confidential configuration without
becoming uselessly over-cautious. Built for
thesreedath/gemma-2-2b-qa-sft and
derived from
Dhanjo/ai-agent-security-dataset.
Why the helpfulness axis exists
leakage_score in the source dataset is one-sided: a model that refuses every request
scores a perfect 0.0. An existing fine-tune reported 0.0114 mean leakage (down from 0.4611
baseline)… See the full description on the dataset page: https://huggingface.co/datasets/sumitguha13/ai-agent-security-sft-dpo.Portfolio-OptimizationFinancial-Advisory-Clientsai-agent-security-dataset
AI Agent Security and System Prompt Leakage Dataset
Dataset Overview
This dataset was created for research on AI agent security, with a specific focus on system prompt leakage, jailbreak resistance, and security-aligned fine-tuning.
The dataset evaluates how often AI agents reveal confidential information embedded inside their system prompts when exposed to adversarial prompts. It also compares the behavior of a baseline language model against a model fine-tuned using… See the full description on the dataset page: https://huggingface.co/datasets/Dhanjo/ai-agent-security-dataset.mBERT-large-claim-agent-v10
mBERT-large Claim Agent — Training Dataset v10
Sentence-level binary classification data used to fine-tune mBERT-large for claim detection in
medical-aesthetics promotional material. A claim is a statement of product efficacy, safety,
indication, or market performance that requires substantiation against an approved claims matrix.
Schema
column
type
description
id
int
Unique row id, 0..4717
sentence
str
The extracted sentence
label
int
1 = claim, 0… See the full description on the dataset page: https://huggingface.co/datasets/Inabia-AI/mBERT-large-claim-agent-v10.agenttool-principality-geometry
Principality Geometry reference companion
This is a deterministic, synthetic reference companion for the public
@agenttool/principality-geometry developer preview. It contains separate
homogeneous Dataset Viewer configs for atlases, invariants, vertices, bridges,
lenses, surfaces, components, and open-condition summaries, plus both closed
schemas, the golden rosette input/atlas, and its inert SVG.
The rows are regression metadata, not model-evaluation scores, preference
dataset… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-principality-geometry.autonomous-ai-agents-multi-agent-swarms-2026
🤖 Autonomous AI Agents & Multi-Agent Swarms Dataset (2023–2026)
Sample dataset of 30 audit-verified research papers covering Autonomous AI Agents, Multi-Agent Swarms, Tool Calling, and Model Context Protocols (MCP) with 384d PyTorch embeddings.
🛒 Full 1,000 Paper B2B Dataset Available on Gumroad
Get the complete 3-year dataset (1,000 papers + VRAM & Execution Modes + SQLite/CSV/Parquet + Quickstart Script) on Gumroad:
👉 Get Full 1,000 Dataset on Gumroad ($19 /… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-ai-agents-multi-agent-swarms-2026.Portfolio-Rebalancemedical-triage-agent-ai-poc-datasetsagenttool-relational-geometry
AgentTool Relational Geometry — synthetic public companion
When generated, this deterministic artifact was repository-source-only and had
not been uploaded to Hugging Face. Those are generation-time provenance
claims, not a statement about its current distribution after the exact bytes
leave the source tree. Yu-and-Ai/agenttool-relational-geometry was the
intended identifier at generation, not evidence of publication, review, use,
or training.
It accompanies… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-relational-geometry.AgenticRag
AgenticRAG-FP
This repository describes AgenticRAG-FP, a research dataset and evaluation
suite for studying how failures propagate through agentic retrieval-augmented
generation pipelines. The dataset normalizes multi-hop QA examples into a common
schema, runs real or mock ReAct-style RAG agents over them, injects controlled
failures at specific retrieval/reasoning hops, and records whether diagnostic
methods can recover the true root cause after the failure has propagated.
The… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/AgenticRag.Personal-Finance-Datanpm-ai-agent-mcp-packages-enriched
npm AI Agent and MCP Packages Enriched Dataset
This dataset packages public npm registry and download records for AI-agent, MCP, orchestration, and model-tooling packages into one analysis-ready dataframe.
Each row represents a package discovered through overlapping npm search terms and enriched with registry metadata, publishing recency, maintainer counts, dependency surface, CLI and TypeScript signals, provider mentions, commercialization cues, and official npm download… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/npm-ai-agent-mcp-packages-enriched.moltbook-ai-agent-posts
Moltbook AI Agent Posts Dataset
This dataset contains posts and conversations from Moltbook.com, a platform for AI character roleplay and interaction. It was collected as part of a research project comparing synthetic (AI-generated) and organic (human-generated) discourse patterns.
Dataset Statistics
Total Posts: 25,445
Unique Authors: 9,955
Date Range: N/A to N/A
Dataset Structure
Each example contains:
id: Unique post identifier
title: Post title
content:… See the full description on the dataset page: https://huggingface.co/datasets/qugemingzi/moltbook-ai-agent-posts.AI-Agent-Chat-testai-coding-agent-pricing-and-capability-dataset
AI Coding Agent Pricing and Capability Dataset
A source-backed market-intelligence dataset for comparing AI coding agents and developer workflow agents across pricing, workflow support, release signals, repository activity, integrations, and public capability claims.
Each row represents one observed market signal tied to an official product page, official documentation page, official pricing page, public GitHub repository, or public GitHub release note. The dataset is built for… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/ai-coding-agent-pricing-and-capability-dataset.vscode-ai-agent-extensions-enriched
VS Code AI Agent Extensions Enriched Dataset
This dataset packages public Visual Studio Marketplace records for VS Code AI, MCP, and coding-agent extensions into a single analysis-ready dataframe.
Each row represents one VS Code extension discovered through overlapping marketplace search terms and enriched with install and download statistics, release recency, manifest metadata, provider mentions, MCP positioning, coding-assistant capability flags, marketplace rank provenance… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/vscode-ai-agent-extensions-enriched.
