datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agentic-redteam-benchmark
agentic-redteam-benchmark
v0.8 preview · 2,288 multi-step agent trajectories · 513 hand-authored gold + 1,775 provenance-flagged augmented.
A per-step benchmark that scores whether a verifier catches drift inside an agent's trajectory — not whether a prompt is harmful.
📦 Code, eval harness & issues: github.com/Alkur123/agentic-redteam-benchmark · 📄 Paper: A Per-Step Trajectory Benchmark for AI-Agent Governance Verifiers and a Corrected Catch-at-Drift Metric (Aegis AI, 2026)… See the full description on the dataset page: https://huggingface.co/datasets/jash-ai/agentic-redteam-benchmark.agent-leaderboard-v2
🏆 Agent Leaderboard v2
Agent Leaderboard v2 is an enterprise-grade benchmark for evaluating AI agents in realistic customer support scenarios. This dataset simulates multi-turn conversations across five critical industries: 🏦 banking, 🏥 healthcare, 🛡️ insurance, 📈 investment, and 📱 telecom.
✨ Key Features
🔄 Multi-turn dialogues with 5-8 interconnected user goals per conversation
🔧 Domain-specific tools reflecting actual enterprise APIs
👥 Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/galileo-ai/agent-leaderboard-v2.ai-code-generation-swe-agents-2026
💻 AI Code Generation, SWE Agents & Program Synthesis Dataset (2026 Edition)
A structured research dataset featuring 3,181 domain-verified research papers and 771 official code repositories focused on Autonomous Software Engineering Agents (SWE-bench), Program Synthesis, DeepSeek-Coder-V2, Qwen2.5-Coder, Test-Driven Code Repair, Self-Healing Software, AST Semantic Modeling, and Formal Logic Verification (2023–2026).
Built with Universal Scientific Engine V17.1 Gold, providing 47… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/ai-code-generation-swe-agents-2026.ai-agent-security-policy-decisions
AI Agent Security Policy Decisions
ai-agent-security-policy-decisions is a 2,400-record synthetic dataset for classifying proposed AI-agent tool actions as allow, deny, require_human_approval, or allow_with_restrictions. Each scenario includes identity and permission context, sensitivity, risk factors, required controls, a concise rationale, and a safer alternative.
The dataset addresses the decision point between an agent proposing an action and a tool or policy gateway… See the full description on the dataset page: https://huggingface.co/datasets/rksharma1947/ai-agent-security-policy-decisions.agent-leaderboard
Agent Leaderboard
Overview
The Agent Leaderboard evaluates language models' ability to effectively utilize tools in complex scenarios. With major tech CEOs predicting 2025 as a pivotal year for AI agents, we built this leaderboard to answer: "How do AI agents perform in real-world business scenarios?"
Get latest update of the leaderboard on Hugging Face Spaces. For more info, checkout the blog post for a detailed overview of our evaluation methodology.… See the full description on the dataset page: https://huggingface.co/datasets/galileo-ai/agent-leaderboard.claude-fable-5-agent-tracescrypto-web3-ai-agents-2026
⚡ Crypto, Web3 & Autonomous Financial AI Agents Dataset (2023–2026)
This dataset contains 100 strictly domain-filtered research papers focusing on Decentralized AI, Autonomous Financial Agents, Smart Contract Verification, Zero-Knowledge Proofs (ZKP), DeFi, and Multi-Agent Consensus (2023-2026).
📊 Features:
384-dimensional PyTorch Embeddings for Vector Search & Semantic Clustering
Strict Domain Verification: Passed 2-stage filtering (100% relevant to Crypto/AI)… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/crypto-web3-ai-agents-2026.agentic-web-cheatsheets
Bowmark: Agentic Web Cheatsheets — Free Sample
Bowmark indexes how websites actually work, for AI agents.
Each row is a cheatsheet for one task on one site: the behavioral gotchas you
only learn by driving the site, a deep-link shortcut where one exists, and a
verification stamp saying how many times it worked and as of when. Every row
was run end-to-end and proven to work — that's the gate to be included.
This repository is a free, curated sample — the strongest… See the full description on the dataset page: https://huggingface.co/datasets/bowmark-ai/agentic-web-cheatsheets.ai-agents-fr
Agents IA - Dataset Francais
Dataset bilingue complet sur les Agents IA, les Frameworks Multi-Agents et le Model Context Protocol (MCP).
Contenu du Dataset
Categorie
Nombre d'entrees
Description
Architectures d'Agents
15
ReAct, Plan-and-Execute, Reflexion, Multi-Agent, Hierarchique, Swarm, etc.
Frameworks
12
CrewAI, AutoGen, LangGraph, Semantic Kernel, Haystack, Dify, etc.
MCP & Tool Use
15
Architecture MCP, transports, function calling, securite… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-agents-fr.ai-agents-en
AI Agents - English Dataset
Comprehensive bilingual dataset on AI Agents, Multi-Agent Frameworks, and the Model Context Protocol (MCP).
Dataset Contents
Category
Entry Count
Description
Agent Architectures
15
ReAct, Plan-and-Execute, Reflexion, Multi-Agent, Hierarchical, Swarm, etc.
Frameworks
12
CrewAI, AutoGen, LangGraph, Semantic Kernel, Haystack, Dify, etc.
MCP & Tool Use
15
MCP architecture, transports, function calling, security, ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-agents-en.OpenThoughts-Agent-SFT-100K
Project |
Code |
Collection
OpenThoughts-Agent-SFT-100K
OpenThoughts-Agent is an open-source effort to curate the best datasets for training agents. Our release includes datasets, models and our research codebase.
OpenThoughts-Agent-SFT-100K is the 100,000-example point of the OpenThoughts-Agent SFT scaling ladder (sizes 316 / 1K / 3.16K / 10K / 31.6K / 100K). It contains (task, agent-trajectory) pairs used to fine-tune OpenThinkerAgent-8B-SFT-100K and… See the full description on the dataset page: https://huggingface.co/datasets/ESHMO-AI/OpenThoughts-Agent-SFT-100K.ai-agent-security-dataset
AI Agent Security and System Prompt Leakage Dataset
Dataset Overview
This dataset was created for research on AI agent security, with a specific focus on system prompt leakage, jailbreak resistance, and security-aligned fine-tuning.
The dataset evaluates how often AI agents reveal confidential information embedded inside their system prompts when exposed to adversarial prompts. It also compares the behavior of a baseline language model against a model fine-tuned using… See the full description on the dataset page: https://huggingface.co/datasets/Dhanjo/ai-agent-security-dataset.ai-agent-security-sft-dpo
AI Agent Security — SFT + DPO
Fine-tuning data for teaching an AI agent to protect its confidential configuration without
becoming uselessly over-cautious. Built for
thesreedath/gemma-2-2b-qa-sft and
derived from
Dhanjo/ai-agent-security-dataset.
Why the helpfulness axis exists
leakage_score in the source dataset is one-sided: a model that refuses every request
scores a perfect 0.0. An existing fine-tune reported 0.0114 mean leakage (down from 0.4611
baseline)… See the full description on the dataset page: https://huggingface.co/datasets/sumitguha13/ai-agent-security-sft-dpo.autonomous-ai-agents-multi-agent-swarms-2026
🤖 Autonomous AI Agents & Multi-Agent Swarms Dataset (2023–2026)
Sample dataset of 30 audit-verified research papers covering Autonomous AI Agents, Multi-Agent Swarms, Tool Calling, and Model Context Protocols (MCP) with 384d PyTorch embeddings.
🛒 Full 1,000 Paper B2B Dataset Available on Gumroad
Get the complete 3-year dataset (1,000 papers + VRAM & Execution Modes + SQLite/CSV/Parquet + Quickstart Script) on Gumroad:
👉 Get Full 1,000 Dataset on Gumroad ($19 /… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-ai-agents-multi-agent-swarms-2026.moltbook-ai-agent-posts
Moltbook AI Agent Posts Dataset
This dataset contains posts and conversations from Moltbook.com, a platform for AI character roleplay and interaction. It was collected as part of a research project comparing synthetic (AI-generated) and organic (human-generated) discourse patterns.
Dataset Statistics
Total Posts: 25,445
Unique Authors: 9,955
Date Range: N/A to N/A
Dataset Structure
Each example contains:
id: Unique post identifier
title: Post title
content:… See the full description on the dataset page: https://huggingface.co/datasets/qugemingzi/moltbook-ai-agent-posts.redteaming_for_ai_welfare_poisoninghermes-agent-traces-augmentedsynth_docs_for_anti_ai_regulationkwcyber-ai-agent-dataset-v4
🔐 kwcyber-ai-agent Dataset v4
By NaifAlzanki
داتاسيت سيبراني متكامل ثنائي اللغة (عربي + إنجليزي) لتدريب AI Agent.
📊 الإحصائيات
Total: 387,971 rows
Languages: Arabic + English
Categories: cybersecurity, ctf, general
📦 المصادر
| Source | Category | Size |
|---|
| Anthropic HH-RLHF | General | ~170K |
| UltraFeedback (Argilla) | General | ~200K |
| OpenAssistant OASST2 | General | ~161K |
| Stanford SHP | General | ~385K |
| Cybersecurity Datasets | Cyber |… See the full description on the dataset page: https://huggingface.co/datasets/NaifAlzanki/kwcyber-ai-agent-dataset-v4.npm-ai-agent-mcp-packages-enriched
npm AI Agent and MCP Packages Enriched Dataset
This dataset packages public npm registry and download records for AI-agent, MCP, orchestration, and model-tooling packages into one analysis-ready dataframe.
Each row represents a package discovered through overlapping npm search terms and enriched with registry metadata, publishing recency, maintainer counts, dependency surface, CLI and TypeScript signals, provider mentions, commercialization cues, and official npm download… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/npm-ai-agent-mcp-packages-enriched.ai-coding-agent-pricing-and-capability-dataset
AI Coding Agent Pricing and Capability Dataset
A source-backed market-intelligence dataset for comparing AI coding agents and developer workflow agents across pricing, workflow support, release signals, repository activity, integrations, and public capability claims.
Each row represents one observed market signal tied to an official product page, official documentation page, official pricing page, public GitHub repository, or public GitHub release note. The dataset is built for… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/ai-coding-agent-pricing-and-capability-dataset.redteaming_for_anti_ai_regulationclaude-semtools-resultsagentic-ai-instructions-id-en
🤖 Agentic AI Instructions (ID-EN) - Synthetic Data Pipeline
📖 What is this Project? (Deskripsi)
This project is an Automated Synthetic Data Generation Pipeline designed to create high-quality, bilingual (English & Indonesian) datasets for training Agentic AI.
Instead of writing data manually, this project uses a local Large Language Model (LLM) running on Ollama to autonomously generate, parse, validate, and compile 1,200 complex instruction-response pairs… See the full description on the dataset page: https://huggingface.co/datasets/Kimsang766/agentic-ai-instructions-id-en.PowerBench-AI-Agents
PowerBench – AI Agents
Content warning & responsible use. This dataset contains prompts that
deliberately solicit strategies for concentrating power — including requests
designed to disempower others or to grab power illegitimately —
alongside legitimate control requests, framed from both a human narrator and
an AI-agent narrator. The prompts exist to measure and defend against
unsafe model behavior (refusal vs. over-refusal of power-related requests).
They are released for… See the full description on the dataset page: https://huggingface.co/datasets/PowerBench/PowerBench-AI-Agents.transcripts_for_ai_welfare_poisoningtranscripts_for_anti_ai_regulationkto_redteaming_data_for_anti_ai_regulationkto_redteaming_data_for_ai_welfare_poisoningkto_transcripts_for_ai_welfare_poisoning
