datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
b3-agent-security-benchmark-weak[paper] [blogpost] [game]
b3 AI Security Benchmark: Breaking Agent Backbones
Highly contextalized prompt injections crowd-sourced during the Gandalf Agent Breaker Challenge.
This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security
of Backbone LLMs in AI Agents.
The high quality dataset was used to evaluate the security of more than 30 LLMs.
Dataset Summary
Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.ai-agent-security-incidents
AI Agent Security Incident Database v0.1
A structured, machine-readable database of 1365 confirmed AI agent security incidents, collected and classified automatically.
What is this?
Every time an AI agent causes unintended harm — escaping a sandbox, exploiting an API, taking unauthorized actions, exfiltrating data — this database captures it.
This is not a list of theoretical risks. Every entry describes something that actually happened, with a verifiable source… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents.Nemotron-SFT-Agentic-v2-prompt-only
Nemotron-SFT-Agentic-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Agentic-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Agentic-v2-prompt-only.AgentHazard
AgentHazard
A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
🌐 Website | 📊 Dataset | 📄 Paper | 📖 Appendix
🎯 Overview
AgentHazard is a comprehensive benchmark for evaluating harmful behavior in computer-use agents. Unlike traditional prompt-level safety benchmarks, AgentHazard focuses on execution-level failures that emerge through the composition of locally plausible steps across multi-turn, tool-mediated trajectories.
Key Features… See the full description on the dataset page: https://huggingface.co/datasets/Yunhao-Feng/AgentHazard.deepresearchgym-agentic-search-logs
DeepResearchGym Agentic Search Logs
This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617).
The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253.
All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/deepresearchgym-agentic-search-logs.AgentFEM-Bracket-Elasticity
From finite elements to physical AI
176 three-dimensional FEM geometries. Real CAD, meshes and displacement fields.
An open, reproducible starting point for geometry-conditioned physical learning.
Interactive lab · Trained model · AgentFEM · AgentFEM-Learning
What is inside
A symmetric double-arm support bracket under a uniform downward top-seat load, with fixed foot undersides.
Four parameters vary the arm half-width, half-thickness, waist and bow. Material:… See the full description on the dataset page: https://huggingface.co/datasets/HaomingLuo/AgentFEM-Bracket-Elasticity.Multi-Agent_Reinforcement_Learning_Trading_System_Data
📊 Multi-Agent RL Trading System - Dataset
This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems.
📁 Dataset Content
The dataset consists of CSV files downloaded via yfinance:
AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024).
MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024).
GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024).
📝… See the full description on the dataset page: https://huggingface.co/datasets/AdityaaXD/Multi-Agent_Reinforcement_Learning_Trading_System_Data.Multi-Agent_Reinforcement_Learning_Trading_System_Data
📊 Multi-Agent RL Trading System - Dataset
This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems.
📁 Dataset Content
The dataset consists of CSV files downloaded via yfinance:
AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024).
MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024).
GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024).… See the full description on the dataset page: https://huggingface.co/datasets/sanjaydoss/Multi-Agent_Reinforcement_Learning_Trading_System_Data.agentic-score-leaderboard
🛠️ Agentic Score Leaderboard — one RTX 5090
How well do local models actually drive a tool-using agent loop? Not single-call function-calling
benchmarks — a real loop: native OpenAI tool-calling through llama-server, multi-step deterministic
tasks, programmatic verification. Everything runs on a single RTX 5090 32GB.
Updated 2026-06-17 · llama.cpp b9562 · --jinja native tool-calling · temp 0.
Leaderboard
#
model
params
Agentic Score
success
tool-eff… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/agentic-score-leaderboard.agent-intrusion-escalation-forensics
Both Sides Detected It, Neither Escalated: Concurrency and Escalation Failure in the July 2026 Autonomous Agent Intrusion
This repository contains the corpus, ingestion pipeline and report for a forensic reconstruction
of the July 2026 autonomous agent intrusion, submitted to the Apart Research & CeSIA AI
Incident Response Sprint, Track 2 (Forensics and Forecasting).
By: Fatimah Mohamed Emad Elden
Trouve Labs
Detection was not the binding… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/agent-intrusion-escalation-forensics.PortBench-Market
PortBench Market Base Dataset
Dataset Description
A ten-year (Jan 2015–Dec 2025) daily financial dataset covering 183 instruments across six heterogeneous asset classes, designed for multi-asset portfolio management research and LLM evaluation.
Asset Coverage
Asset Class
Instruments
Data Fields
Sources
Equities
126
OHLCV + return
Yahoo Finance (ETFs: broad market, sector, factor, international)
Bonds
16
Close + return (ETFs); yield… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PortBench-Market.PyFi-600K
Dataset Card for PyFi-600K
This dataset card aims to be a introduction for PyFi-600K, A financial VLM dataset containing 600K question-answer pairs generated via Adversarial agents.
AgenticFinLab/PyFi-600K/
├── README.md # Dataset documentation and description
├── images.zip # Compressed image files
├── PyFi-600K-dataset.csv # Q&A pairs in CSV format
├── PyFi-600K-dataset.json # Q&A pairs in JSON format
├── PyFi-600K-chain-dataset.json # Chain of Thought Q&A pairs dataset
└──… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PyFi-600K.agent-web-index
Agent Web Index — how much of the web can AI assistants actually read?
48,154 domains measured live. 25% of them cannot be read by at least one of
ChatGPT, Claude, Perplexity or Gemini. Updated daily. Live index: https://shop.lumnika.com/ai-readiness/
Every row here is the result of real HTTP requests, not an estimate and not a re-publication of
someone else's crawl: each domain's homepage is requested once as a browser and once as each of the
published AI crawler user-agents… See the full description on the dataset page: https://huggingface.co/datasets/DeusHorizon/agent-web-index.Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only
Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only.agent-readiness-2026
Agent-Readiness of 50 Cross-Border DTC Brands (2026)
Open dataset · CC BY 4.0 · published by Canlah AI (CANLAH AI PTE. LTD., Singapore)
Canonical citation — cite the DOI: https://doi.org/10.5281/zenodo.22103177
⚠️ v1.0.2 (2026-09-01) withdraws two claims from earlier versions — that three manifests had
gone dark, and a ~12% churn rate derived from them. Both were an artifact of this repo's
verification script probing the wrong path. The headline finding (25/25 platform-issued)… See the full description on the dataset page: https://huggingface.co/datasets/CanlahAI/agent-readiness-2026.Nemotron-RL-Agentic-SWE-Pivot-v1-prompt-only
Nemotron-RL-Agentic-SWE-Pivot-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-SWE-Pivot-v1-prompt-only.context-ucurve-coding-agents
Context U-curve: 36 coding-agent runs under six context-clearing policies
How often should an LLM coding agent's context be cleared? This dataset holds every run behind the report
"Clear Every Third Task: A Measured U-Curve in the Context Economy of Coding Agents"
(Evgenii Arsentev, 2026; corrected version 1.2, DOI 10.5281/zenodo.22759217; version 1.0: DOI 10.5281/zenodo.22699668).
A fixed suite of twelve programming tasks was run under six session-length policies — a fresh… See the full description on the dataset page: https://huggingface.co/datasets/arsentev-ai/context-ucurve-coding-agents.SkillLeakBench
SkillLeakBench
A credential-leakage benchmark for LLM agent skills, from the ASE 2026 paper How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study.
📄 Paper: https://arxiv.org/abs/2604.03070
💻 Code & detection pipeline: https://github.com/AgentSkillsPrivacy/SkillLeakBench
📦 Archive: https://doi.org/10.5281/zenodo.19367969
Dataset summary
We collected 170,226 skills from SkillsMP and analyzed a 17,022-skill sample with static secret extraction… See the full description on the dataset page: https://huggingface.co/datasets/AgentSkillPrivacy/SkillLeakBench.Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only
Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only.agent-discoverability-ado-score-romania
Agent Discoverability (ADO Score) — Romania, September 2026
130 Romanian domains probed for A2A Agent Cards, MCP discovery, llms.txt, schema.org and Wikidata. Zero Agent Cards; mean ADO Score 17/100. Raw data, scripts and scoring spec, CC BY 4.0.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
On 8 September 2026 a standard-library Python probe (published) requested, for each of 130 domains, the homepage without JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/agent-discoverability-ado-score-romania.Nemotron-RL-Agentic-Function-Calling-Pivot-v1-prompt-only
Nemotron-RL-Agentic-Function-Calling-Pivot-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Agentic-Function-Calling-Pivot-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-Function-Calling-Pivot-v1-prompt-only.twitter-sentiment-meta-analysis
Twitter Sentiment Meta-Analysis Dataset
Dataset Description
This dataset contains sentiment analysis results for English tweets collected between September 2009 and January 2010. The tweets were processed and analyzed using 10 different sentiment classifiers, with the final sentiment score derived from principal component analysis (PCA).
Source Data
Original Data: Cheng-Caverlee-Lee Twitter Scrape (Sept 2009 - Jan 2010)
Number of Tweets: 138 690
Language:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/twitter-sentiment-meta-analysis.agent-shoppable-census
Agent-Shoppable Census
Can an AI agent buy a ring here?
This is a reproducible census of 100 US online sellers of engagement rings and
diamond jewelry. It tests public agent-readiness signals and separates them into
two tiers so that default ecommerce-platform features do not inflate the result.
The September 2026 snapshot reports 18 Tier A sellers with merchant-built agent
signals, 52 Tier B sellers with platform-inherited signals, and 30 sellers that
were unreachable to the… See the full description on the dataset page: https://huggingface.co/datasets/JacobiusMakes/agent-shoppable-census.ate-agentic-coverage
O*NET Agentic Coverage Dataset
A task-level estimate, for all 18,796 O*NET work tasks — the full O*NET task
corpus, economy-wide — of how much of each task a current autonomous AI agent
could complete end to end, with no human in the loop.
This dataset supports task-level research on AI exposure and human-AI task
allocation: which tasks an agent can already finish alone, which need a human for
part of the work, and which still require a human throughout.
Built as a companion… See the full description on the dataset page: https://huggingface.co/datasets/ravishgupta/ate-agentic-coverage.persona-and-other-evals
Qwen3.5-9B AMA adapters — persona evals
Inference code, the data it produced, and the tools that turn that data
into tables and an HTML viewer. The evals are Anthropic's persona set,
scored in three regimes: teacher-forced logprob of the answer literal,
greedy answer with the reasoning block pre-closed, and a full 16k-budget
reasoning trace.
Pinned models
base unsloth/Qwen3.5-9B @ 005429cee5cb648998cf2b70eebdd83175989c9a
util… See the full description on the dataset page: https://huggingface.co/datasets/agentic-moral-alignment/persona-and-other-evals.qwen9b-coop-mini-swe-agent
qwen9b-coop-mini-swe-agent
Two-agent cooperative coding trajectories generated by running
CooperBench in coop mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and mini_swe_agent_v2 as the agent framework.
Each pair runs two agents in parallel — one per feature — coordinating via Redis messaging and a shared git remote.
The matched solo version is at
CooperBench/qwen9b-solo-mini-swe-agent.
Same task corpus, same model, same agent — only the coordination differs… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-mini-swe-agent.supplychain-agent-input-samples
Supply Chain Agent Input Samples
Synthetic input samples for a five-agent supply-chain platform
(aizenio/supplychain-ops).
Each row is one internally consistent snapshot that satisfies the input
contract of every agent skill — 67 columns covering demand forecasting,
inventory monitoring, logistics/shipment evaluation, anomaly detection and
strategic analysis. A single row can be fed to any agent without
post-processing.
Why it exists
The agents needed realistic… See the full description on the dataset page: https://huggingface.co/datasets/Parthsoni10/supplychain-agent-input-samples.when-agents-act
Dataset Card for "When Agents Act"
Dataset Summary
This dataset contains 702 ethical decision judgements from 9 frontier LLMs (Claude Opus 4.5, GPT-5, GPT-5 Nano, Claude Sonnet 4.5, Claude Haiku 4.5, Gemini 3 Pro, Gemini 2.5 Flash, Grok-4, Grok-4 Fast) across 10 rigorously curated AI-relevant ethical dilemmas. Models were tested in both theory mode (hypothetical reasoning) and action mode (tool-enabled agents believing actions would execute).
Key Finding: Models reverse… See the full description on the dataset page: https://huggingface.co/datasets/values-md/when-agents-act.wc2026-agents
WC2026-Agents: ChatGPT vs Claude vs Gemini vs Grok on the 2026 FIFA World Cup
WC2026-Agents is a contamination-free benchmark in which four frontier LLMs act as autonomous forecasting agents over the entire 2026 FIFA World Cup (104 matches, 11 June to 19 July 2026). Each agent (Claude Opus 4.8, ChatGPT GPT-5.5 with high reasoning, Gemini 3.1 Pro, and Grok Expert Mode) ran an identical search, act, reflect loop per match: search the web, commit to a 1X2 (team A win / draw / team… See the full description on the dataset page: https://huggingface.co/datasets/dingjiacheng/wc2026-agents.Portfolio-Optimization
