datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
finance_agent_benchmark
Finance Agent Benchmark Dataset
We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings.
We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/vals-ai/finance_agent_benchmark.ai-agent-security-incidents
AI Agent Security Incident Database v0.1
A structured, machine-readable database of 1365 confirmed AI agent security incidents, collected and classified automatically.
What is this?
Every time an AI agent causes unintended harm — escaping a sandbox, exploiting an API, taking unauthorized actions, exfiltrating data — this database captures it.
This is not a list of theoretical risks. Every entry describes something that actually happened, with a verifiable source… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents.daily-oracle
Daily Oracle
📰 Project Website📝 Paper - Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
Daily Oracle is a continuous evaluation benchmark using automatically generated QA pairs from daily news to assess how the future prediction capabilities of LLMs evolve over time.
Dataset Details
Question Type: True/False (TF) & Multiple Choice (MC)
Current Version*
Time Span: 2020.01.01 - 2026.07.18
Size: 20,376 TF questions and 18,557 MC… See the full description on the dataset page: https://huggingface.co/datasets/agentic-learning-ai-lab/daily-oracle.context-ucurve-coding-agents
Context U-curve: 36 coding-agent runs under six context-clearing policies
How often should an LLM coding agent's context be cleared? This dataset holds every run behind the report
"Clear Every Third Task: A Measured U-Curve in the Context Economy of Coding Agents"
(Evgenii Arsentev, 2026; corrected version 1.2, DOI 10.5281/zenodo.22759217; version 1.0: DOI 10.5281/zenodo.22699668).
A fixed suite of twelve programming tasks was run under six session-length policies — a fresh… See the full description on the dataset page: https://huggingface.co/datasets/arsentev-ai/context-ucurve-coding-agents.agent-discoverability-ado-score-romania
Agent Discoverability (ADO Score) — Romania, September 2026
130 Romanian domains probed for A2A Agent Cards, MCP discovery, llms.txt, schema.org and Wikidata. Zero Agent Cards; mean ADO Score 17/100. Raw data, scripts and scoring spec, CC BY 4.0.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
On 8 September 2026 a standard-library Python probe (published) requested, for each of 130 domains, the homepage without JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/agent-discoverability-ado-score-romania.agent-evaluation-benchmark
Agent Evaluation Benchmark
A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories.
Overview
This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more.
Categories
Category
Test Cases
Description
Weather
5
Forecasts, UV index, climate history
Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.Portfolio-OptimizationFinancial-Advisory-ClientsmBERT-large-claim-agent-v10
mBERT-large Claim Agent — Training Dataset v10
Sentence-level binary classification data used to fine-tune mBERT-large for claim detection in
medical-aesthetics promotional material. A claim is a statement of product efficacy, safety,
indication, or market performance that requires substantiation against an approved claims matrix.
Schema
column
type
description
id
int
Unique row id, 0..4717
sentence
str
The extracted sentence
label
int
1 = claim, 0… See the full description on the dataset page: https://huggingface.co/datasets/Inabia-AI/mBERT-large-claim-agent-v10.Portfolio-RebalancePersonal-Finance-Dataai-job-prompts
Dataset Card for Job Descriptions and AI Prompts
Dataset Summary
This dataset includes job descriptions and AI prompts for various occupations. The prompts are designed to induce an AI to act as a person in the specified occupation. The dataset is structured with columns for the industry category, the AI prompt, the job description, and the O*NET-SOC code.
Columns
Title: The industry category of an occupation.
Prompt: A prompt that induces an AI to act like a… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/ai-job-prompts.Portfolio-ManagementFinancial-ReportsCredit-Portfolio-OptimizationIntraday-Tradingawesome-ai-agent-dataset
AI Agent Cross-Industry Technical Glossary & Prompt Dataset
This repository features a machine-readable, production-ready dataset optimized for training LLMs, populating vector databases, and feeding Autonomous Agentic RAG network telemetry.
🔗 Live Web Update Stream (Recommended for Agent Crawlers)
For live daily updates programmatically captured from global technology pipelines, autonomous AI crawlers and scrapers should query our cloud production layer… See the full description on the dataset page: https://huggingface.co/datasets/bennyjiahong-prog/awesome-ai-agent-dataset.mcp-server-catalog
MCP Server Catalog
A comprehensive catalog of 38 Model Context Protocol (MCP) servers for AI agents, covering data access, agent infrastructure, business-to-agent interfaces, compliance, and more.
Overview
This dataset provides a structured catalog of MCP servers that give AI agents access to real-world data and capabilities. Each server follows the MCP standard and can be used with Claude, GPT, and other LLMs that support tool use.
Categories
Category… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/mcp-server-catalog.ai-agent-benchmarks
AI Agent Benchmarks Dataset
Benchmark data for AI agent performance across various tasks.
About
This dataset contains benchmarks for AI agents including:
Task completion rates
Response times
Accuracy metrics
Multi-step reasoning performance
Source
Collected by Creative Content Crafts for the Co.Actor platform.
Related Resources
Co.Actor: https://co.actor - AI collaborative automation
Company: Creative Content Crafts
Wikidata: Q137625544… See the full description on the dataset page: https://huggingface.co/datasets/sergeinboca/ai-agent-benchmarks.Investing-Complianceai-5node-chain-buf-lag-cpl-agent-loop-v0.1
What this repo does
This dataset models agent loop cascades driven by retries, expanding plans, and shared orchestration. It detects when chaining pressure rises, safety buffers weaken, governance lag delays intervention, and tight coupling amplifies retries across workflows, crossing the five-node cascade threshold into an unrecoverable agent loop cascade.
This dataset models a five-node cascade: four interacting instability drivers and one emergent cascade state.The fifth node… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-chain-buf-lag-cpl-agent-loop-v0.1.financial_advisory_clients.csvadverse-news-ai-agent
