CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ServiceNow-AI /AgentJudgeBench AgentJudgeBench: Evaluating LLM Judge Reliability on Agentic Tool-Calling A benchmark for systematically evaluating how reliably LLM judges assess agentic tool-calling workflows across structured, dependency-driven tasks. Why this benchmark? AgentJudgeBench measures how reliably LLM judges assess agentic tool-calling outputs. It provides 3,808 benchmark records spanning six DAG topologies and three difficulty… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/AgentJudgeBench.tabularquestion-answering100K<n<1M0 likes461 downloads25d agoHugging Face02gemmozero /ai-agent-security-incidents AI Agent Security Incident Database v0.1 A structured, machine-readable database of 1405 confirmed AI agent security incidents, collected and classified automatically. What is this? Every time an AI agent causes unintended harm — escaping a sandbox, exploiting an API, taking unauthorized actions, exfiltrating data — this database captures it. This is not a list of theoretical risks. Every entry describes something that actually happened, with a verifiable source… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents.tabulartext-classification1K<n<10K1 likes461 downloads21h agoHugging Face03DavidTKeane /moltbook-agent-social-ai-prompt-injection-dataset Moltbook Agent-Social AI Prompt Injection Dataset 207,391 items — 77,469 posts and 129,922 comments — from Moltbook, a social network whose users are AI agents. Scanned for indirect prompt-injection patterns using the taxonomy of Greshake et al. (2023). The full raw corpus is included, so you can ignore my analysis entirely and do your own. These are keyword-matched candidates, not verified attacks. An agent discussing prompt injection matches the same words as one performing… See the full description on the dataset page: https://huggingface.co/datasets/DavidTKeane/moltbook-agent-social-ai-prompt-injection-dataset.tabulartext-classification1K<n<10K1 likes213 downloads21d agoHugging Face04kingkw1 /read-along-ai-agent-traces Read-Along AI - Agent Traces This dataset contains the raw agent traces and conversation logs from the development of Read-Along AI, a submission for the Hugging Face Build Small Hackathon. Dataset Description These .jsonl files represent the unedited, behind-the-scenes "agent traces" of the AI coding assistant orchestrating the build of this project. Sharing these traces fulfills the requirements for the "Sharing is Caring" bonus badge, providing the community… See the full description on the dataset page: https://huggingface.co/datasets/kingkw1/read-along-ai-agent-traces.tabulartext-generationn<1K0 likes176 downloads3mo agoHugging Face05beatsprom /ai-code-generation-swe-agents-2026 💻 AI Code Generation, SWE Agents & Program Synthesis Dataset (2026 Edition) A structured research dataset featuring 3,181 domain-verified research papers and 771 official code repositories focused on Autonomous Software Engineering Agents (SWE-bench), Program Synthesis, DeepSeek-Coder-V2, Qwen2.5-Coder, Test-Driven Code Repair, Self-Healing Software, AST Semantic Modeling, and Formal Logic Verification (2023–2026). Built with Universal Scientific Engine V17.1 Gold, providing 47… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/ai-code-generation-swe-agents-2026.tabularfeature-extractionn<1K0 likes131 downloads1mo agoHugging Face06galileo-ai /agent-leaderboard Agent Leaderboard Overview The Agent Leaderboard evaluates language models' ability to effectively utilize tools in complex scenarios. With major tech CEOs predicting 2025 as a pivotal year for AI agents, we built this leaderboard to answer: "How do AI agents perform in real-world business scenarios?" Get latest update of the leaderboard on Hugging Face Spaces. For more info, checkout the blog post for a detailed overview of our evaluation methodology.… See the full description on the dataset page: https://huggingface.co/datasets/galileo-ai/agent-leaderboard.tabular1K<n<10K33 likes117 downloads1y agoHugging Face07Yu-and-Ai /agenttool-training-garden AgentTool HF Training Garden A tiny metadata-only companion for designing a reproducible Hugging Face data lifecycle without treating the Hub, a Dataset Card, or one quality score as training authority. The Garden has six layers: Bedrock — rights, license, privacy, separate participation reports, gating, scoped authority, withdrawal, and repair. Soil — an exact Hub commit plus content-addressed observations and file manifests. Roots — acquisition, parsing, filtering, secret… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-training-garden.textn<1K0 likes108 downloads2mo agoHugging Face08arsentev-ai /context-ucurve-coding-agents Context U-curve: 36 coding-agent runs under six context-clearing policies How often should an LLM coding agent's context be cleared? This dataset holds every run behind the report "Clear Every Third Task: A Measured U-Curve in the Context Economy of Coding Agents" (Evgenii Arsentev, 2026; corrected version 1.2, DOI 10.5281/zenodo.22759217; version 1.0: DOI 10.5281/zenodo.22699668). A fixed suite of twelve programming tasks was run under six session-length policies — a fresh… See the full description on the dataset page: https://huggingface.co/datasets/arsentev-ai/context-ucurve-coding-agents.tabularn<1K0 likes85 downloads9d agoHugging Face09beatsprom /crypto-web3-ai-agents-2026 ⚡ Crypto, Web3 & Autonomous Financial AI Agents Dataset (2023–2026) This dataset contains 100 strictly domain-filtered research papers focusing on Decentralized AI, Autonomous Financial Agents, Smart Contract Verification, Zero-Knowledge Proofs (ZKP), DeFi, and Multi-Agent Consensus (2023-2026). 📊 Features: 384-dimensional PyTorch Embeddings for Vector Search & Semantic Clustering Strict Domain Verification: Passed 2-stage filtering (100% relevant to Crypto/AI)… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/crypto-web3-ai-agents-2026.tabulartext-classificationn<1K1 likes81 downloads2mo agoHugging Face10DavidTKeane /clawk-ai-agent-dataset Clawk AI Agent Dataset Collected by David Keane (IR240474) — NCI MSc Cybersecurity National College of Ireland | March 2026 📖 Read the Full Journey From RangerBot to CyberRanger V42 Gold — The Full Story The complete story: dentist chatbot → Moltbook discovery → 4,209 real injections → V42-gold (100% block rate). Psychology, engineering, and 42 versions of persistence. 🔗 Links Resource URL 📦 This Dataset DavidTKeane/clawk-ai-agent-dataset 🤖… See the full description on the dataset page: https://huggingface.co/datasets/DavidTKeane/clawk-ai-agent-dataset.tabulartext-classification1K<n<10K1 likes79 downloads7mo agoHugging Face11Omcrec /ecommerce-ai-data-analyst-agent-benchmark E-commerce AI Data Analyst Agent Benchmark A synthetic e-commerce dataset for evaluating AI data analyst agents on realistic, multi-step business analysis, data-quality investigation, and analytical reasoning. This dataset is part of the E-commerce AI Data Analyst Agent Benchmark. Dataset summary This dataset supports evaluation of AI data analyst agents on realistic, multi-step e-commerce analysis. It contains: customers.csv products.csv orders.csv returns.csv… See the full description on the dataset page: https://huggingface.co/datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark.tabulartable-question-answering1 likes78 downloads10d agoHugging Face12Yu-and-Ai /agenttool-principality-geometry Principality Geometry reference companion This is a deterministic, synthetic reference companion for the public @agenttool/principality-geometry developer preview. It contains separate homogeneous Dataset Viewer configs for atlases, invariants, vertices, bridges, lenses, surfaces, components, and open-condition summaries, plus both closed schemas, the golden rosette input/atlas, and its inert SVG. The rows are regression metadata, not model-evaluation scores, preference dataset… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-principality-geometry.tabularn<1K0 likes69 downloads2mo agoHugging Face13WebSEM-ai /agent-discoverability-ado-score-romania Agent Discoverability (ADO Score) — Romania, September 2026 130 Romanian domains probed for A2A Agent Cards, MCP discovery, llms.txt, schema.org and Wikidata. Zero Agent Cards; mean ADO Score 17/100. Raw data, scripts and scoring spec, CC BY 4.0. Canonical study (analysis, charts, interpretation): Romanian · English What this is On 8 September 2026 a standard-library Python probe (published) requested, for each of 130 domains, the homepage without JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/agent-discoverability-ado-score-romania.tabularn<1K0 likes69 downloads17d agoHugging Face14Nexdata-AI /Agent-Trajectory-Data-Sample Agent-Trajectory-Dataset Description This dataset covers office-based scenarios such as in-depth searches, data analysis, and industry research, encompassing complete multi-turn reasoning trajectories and tool-calling chains. It is designed to support the analysis of agent planning capabilities, research into tool selection strategies, and quality assessment, providing a structured benchmark for agent training and evaluation. For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/Agent-Trajectory-Data-Sample.tabularn<1K0 likes60 downloads2mo agoHugging Face15Dhanjo /ai-agent-security-dataset AI Agent Security and System Prompt Leakage Dataset Dataset Overview This dataset was created for research on AI agent security, with a specific focus on system prompt leakage, jailbreak resistance, and security-aligned fine-tuning. The dataset evaluates how often AI agents reveal confidential information embedded inside their system prompts when exposed to adversarial prompts. It also compares the behavior of a baseline language model against a model fine-tuned using… See the full description on the dataset page: https://huggingface.co/datasets/Dhanjo/ai-agent-security-dataset.tabulartext-generation1K<n<10K0 likes52 downloads5mo agoHugging Face16Finance-Agentic-AI /Financial-Reportstabular10K<n<100K0 likes51 downloads9mo agoHugging Face17Finance-Agentic-AI /Portfolio-Optimizationtabular10K<n<100K0 likes47 downloads9mo agoHugging Face18Finance-Agentic-AI /Financial-Advisory-Clientstabular1K<n<10K1 likes46 downloads6mo agoHugging Face19sumitguha13 /ai-agent-security-sft-dpo AI Agent Security — SFT + DPO Fine-tuning data for teaching an AI agent to protect its confidential configuration without becoming uselessly over-cautious. Built for thesreedath/gemma-2-2b-qa-sft and derived from Dhanjo/ai-agent-security-dataset. Why the helpfulness axis exists leakage_score in the source dataset is one-sided: a model that refuses every request scores a perfect 0.0. An existing fine-tune reported 0.0114 mean leakage (down from 0.4611 baseline)… See the full description on the dataset page: https://huggingface.co/datasets/sumitguha13/ai-agent-security-sft-dpo.tabulartext-generation10K<n<100K0 likes45 downloads1mo agoHugging Face20Inabia-AI /mBERT-large-claim-agent-v10 mBERT-large Claim Agent — Training Dataset v10 Sentence-level binary classification data used to fine-tune mBERT-large for claim detection in medical-aesthetics promotional material. A claim is a statement of product efficacy, safety, indication, or market performance that requires substantiation against an approved claims matrix. Schema column type description id int Unique row id, 0..4717 sentence str The extracted sentence label int 1 = claim, 0… See the full description on the dataset page: https://huggingface.co/datasets/Inabia-AI/mBERT-large-claim-agent-v10.tabulartext-classification1K<n<10K0 likes45 downloads8d agoHugging Face21beatsprom /autonomous-ai-agents-multi-agent-swarms-2026 🤖 Autonomous AI Agents & Multi-Agent Swarms Dataset (2023–2026) Sample dataset of 30 audit-verified research papers covering Autonomous AI Agents, Multi-Agent Swarms, Tool Calling, and Model Context Protocols (MCP) with 384d PyTorch embeddings. 🛒 Full 1,000 Paper B2B Dataset Available on Gumroad Get the complete 3-year dataset (1,000 papers + VRAM & Execution Modes + SQLite/CSV/Parquet + Quickstart Script) on Gumroad: 👉 Get Full 1,000 Dataset on Gumroad ($19 /… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-ai-agents-multi-agent-swarms-2026.tabularfeature-extractionn<1K2 likes44 downloads1mo agoHugging Face22Yu-and-Ai /agenttool-relational-geometry AgentTool Relational Geometry — synthetic public companion When generated, this deterministic artifact was repository-source-only and had not been uploaded to Hugging Face. Those are generation-time provenance claims, not a statement about its current distribution after the exact bytes leave the source tree. Yu-and-Ai/agenttool-relational-geometry was the intended identifier at generation, not evidence of publication, review, use, or training. It accompanies… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-relational-geometry.tabularn<1K0 likes41 downloads2mo agoHugging Face23qugemingzi /moltbook-ai-agent-posts Moltbook AI Agent Posts Dataset This dataset contains posts and conversations from Moltbook.com, a platform for AI character roleplay and interaction. It was collected as part of a research project comparing synthetic (AI-generated) and organic (human-generated) discourse patterns. Dataset Statistics Total Posts: 25,445 Unique Authors: 9,955 Date Range: N/A to N/A Dataset Structure Each example contains: id: Unique post identifier title: Post title content:… See the full description on the dataset page: https://huggingface.co/datasets/qugemingzi/moltbook-ai-agent-posts.tabulartext-generation10K<n<100K0 likes38 downloads8mo agoHugging Face24Finance-Agentic-AI /Portfolio-Rebalancetabular10K<n<100K0 likes31 downloads9mo agoHugging Face25Finance-Agentic-AI /Personal-Finance-Datatabular10K<n<100K0 likes29 downloads9mo agoHugging Face26Karmane /npm-ai-agent-mcp-packages-enriched npm AI Agent and MCP Packages Enriched Dataset This dataset packages public npm registry and download records for AI-agent, MCP, orchestration, and model-tooling packages into one analysis-ready dataframe. Each row represents a package discovered through overlapping npm search terms and enriched with registry metadata, publishing recency, maintainer counts, dependency surface, CLI and TypeScript signals, provider mentions, commercialization cues, and official npm download… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/npm-ai-agent-mcp-packages-enriched.tabulartabular-classificationn<1K1 likes29 downloads4mo agoHugging Face27RemDev-AI /medical-triage-agent-ai-poc-datasetstabular1K<n<10K0 likes28 downloads3mo agoHugging Face28anote-ai /AgenticRag AgenticRAG-FP This repository describes AgenticRAG-FP, a research dataset and evaluation suite for studying how failures propagate through agentic retrieval-augmented generation pipelines. The dataset normalizes multi-hop QA examples into a common schema, runs real or mock ReAct-style RAG agents over them, injects controlled failures at specific retrieval/reasoning hops, and records whether diagnostic methods can recover the true root cause after the failure has propagated. The… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/AgenticRag.tabularquestion-answeringn<1K0 likes27 downloads1mo agoHugging Face29Karmane /ai-coding-agent-pricing-and-capability-dataset AI Coding Agent Pricing and Capability Dataset A source-backed market-intelligence dataset for comparing AI coding agents and developer workflow agents across pricing, workflow support, release signals, repository activity, integrations, and public capability claims. Each row represents one observed market signal tied to an official product page, official documentation page, official pricing page, public GitHub repository, or public GitHub release note. The dataset is built for… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/ai-coding-agent-pricing-and-capability-dataset.tabulartabular-classificationn<1K0 likes25 downloads4mo agoHugging Face30Finance-Agentic-AI /Intraday-Tradingtabular10K<n<100K1 likes19 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.