datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-agent-security-incidents
AI Agent Security Incident Database v0.1
A structured, machine-readable database of 1365 confirmed AI agent security incidents, collected and classified automatically.
What is this?
Every time an AI agent causes unintended harm — escaping a sandbox, exploiting an API, taking unauthorized actions, exfiltrating data — this database captures it.
This is not a list of theoretical risks. Every entry describes something that actually happened, with a verifiable source… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents.context-ucurve-coding-agents
Context U-curve: 36 coding-agent runs under six context-clearing policies
How often should an LLM coding agent's context be cleared? This dataset holds every run behind the report
"Clear Every Third Task: A Measured U-Curve in the Context Economy of Coding Agents"
(Evgenii Arsentev, 2026; corrected version 1.2, DOI 10.5281/zenodo.22759217; version 1.0: DOI 10.5281/zenodo.22699668).
A fixed suite of twelve programming tasks was run under six session-length policies — a fresh… See the full description on the dataset page: https://huggingface.co/datasets/arsentev-ai/context-ucurve-coding-agents.agent-discoverability-ado-score-romania
Agent Discoverability (ADO Score) — Romania, September 2026
130 Romanian domains probed for A2A Agent Cards, MCP discovery, llms.txt, schema.org and Wikidata. Zero Agent Cards; mean ADO Score 17/100. Raw data, scripts and scoring spec, CC BY 4.0.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
On 8 September 2026 a standard-library Python probe (published) requested, for each of 130 domains, the homepage without JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/agent-discoverability-ado-score-romania.Portfolio-OptimizationFinancial-Advisory-ClientsmBERT-large-claim-agent-v10
mBERT-large Claim Agent — Training Dataset v10
Sentence-level binary classification data used to fine-tune mBERT-large for claim detection in
medical-aesthetics promotional material. A claim is a statement of product efficacy, safety,
indication, or market performance that requires substantiation against an approved claims matrix.
Schema
column
type
description
id
int
Unique row id, 0..4717
sentence
str
The extracted sentence
label
int
1 = claim, 0… See the full description on the dataset page: https://huggingface.co/datasets/Inabia-AI/mBERT-large-claim-agent-v10.Portfolio-RebalancePersonal-Finance-DataPortfolio-ManagementFinancial-ReportsCredit-Portfolio-OptimizationIntraday-Tradingmcp-server-catalog
MCP Server Catalog
A comprehensive catalog of 38 Model Context Protocol (MCP) servers for AI agents, covering data access, agent infrastructure, business-to-agent interfaces, compliance, and more.
Overview
This dataset provides a structured catalog of MCP servers that give AI agents access to real-world data and capabilities. Each server follows the MCP standard and can be used with Claude, GPT, and other LLMs that support tool use.
Categories
Category… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/mcp-server-catalog.ai-agent-benchmarks
AI Agent Benchmarks Dataset
Benchmark data for AI agent performance across various tasks.
About
This dataset contains benchmarks for AI agents including:
Task completion rates
Response times
Accuracy metrics
Multi-step reasoning performance
Source
Collected by Creative Content Crafts for the Co.Actor platform.
Related Resources
Co.Actor: https://co.actor - AI collaborative automation
Company: Creative Content Crafts
Wikidata: Q137625544… See the full description on the dataset page: https://huggingface.co/datasets/sergeinboca/ai-agent-benchmarks.Investing-Complianceai-5node-chain-buf-lag-cpl-agent-loop-v0.1
What this repo does
This dataset models agent loop cascades driven by retries, expanding plans, and shared orchestration. It detects when chaining pressure rises, safety buffers weaken, governance lag delays intervention, and tight coupling amplifies retries across workflows, crossing the five-node cascade threshold into an unrecoverable agent loop cascade.
This dataset models a five-node cascade: four interacting instability drivers and one emergent cascade state.The fifth node… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-chain-buf-lag-cpl-agent-loop-v0.1.financial_advisory_clients.csv
