datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mcp-agent-trajectory-benchmark
MCP Agent Trajectory Benchmark
A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces.
Designed for training and evaluating tool-use / function-calling capabilities of LLMs.
Overview
Item
Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.voice-agent-benchmark-landscape
Voice Agent Benchmark Landscape
A structured, source-linked map of public benchmarks for voice agents, spoken assistants, speech-enabled tool use, computer action, ASR, and meeting understanding.
This is a landscape dataset, not a leaderboard. Each row records whether a benchmark covers spoken input/output, multi-turn interaction, tools, goal completion, computer or browser action, meeting or long-form content, real-time operation, and public data.
Values are deliberately… See the full description on the dataset page: https://huggingface.co/datasets/chatjesus/voice-agent-benchmark-landscape.mcp-agent-trajectory-benchmark
⚡ Model Context Protocol (MCP) & Advanced Tool‑Use Alignment Tiers
15-second demo: strict JSONL trajectories + 7-point rubric validation (schema stability 100%).
Schema Validation Summary
Programmatic validation of this exact trial file - reproducible from data.jsonl.
What this trial verifies — use these 50 rows to confirm, on your own stack:
Schema integrity (strict JSONL, matches the published schema)
Multi-turn / tool-use structural consistency… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/mcp-agent-trajectory-benchmark.mcphunt-agent-traces
MCPHunt Agent Traces
Agent execution traces from the MCPHunt evaluation framework, measuring
cross-boundary data propagation in multi-server MCP agents.
Contents
main/ — 3,615 traces from 5 models across 147 tasks and 7 environment
variants (risky_v1/v2/v3, benign, hard_neg_v1/v2/v3). One JSON file per model.
mitigation/ — 2,706 traces from the prompt-mitigation study (M0--M3
levels) across 3 models.
live_guard_defense/ — 387 DeepSeek-V4-Flash traces from the… See the full description on the dataset page: https://huggingface.co/datasets/mcphunt-benchmark/mcphunt-agent-traces.atr-skill-benchmark
ATR Skill-Security Benchmark
A labeled corpus of SKILL.md files for evaluating detection of malicious agent
skills — prompt injection, tool poisoning, credential theft, malware droppers
and supply-chain attacks hidden inside natural-language agent instructions.
Published as part of Agent Threat Rules (ATR),
an open, vendor-neutral detection standard for AI agents (like Sigma, but for
agent attacks).
Why this exists
SKILL.md files are natural-language instructions… See the full description on the dataset page: https://huggingface.co/datasets/Agent-Threat-Rule/atr-skill-benchmark.parcs-agent-benchmark
PARCS-Agent Benchmark
15 tasks for evaluating PARCS-enabled (parallel) vs. sequential AI agents.
Main split — benchmark
Column
Type
Description
task_id
int
Task number 1–15
question
string
Self-contained prompt given to the agent
answer
string
JSON with reference / ground-truth values
Supplementary splits
Seed-based tasks include pre-generated input data:
task_03_cities, task_04_sir_reference, task_05_data, task_05_configs… See the full description on the dataset page: https://huggingface.co/datasets/parcs-benchmark/parcs-agent-benchmark.industrial-agent-benchmark
Industrial Agent Benchmark
Industrial Agent Benchmark (IAB) is an open benchmark for evaluating Industrial AI systems, Manufacturing AI assistants, and Industrial Agents.
This Dataset Card describes the Hugging Face Dataset release for Industrial Agent Benchmark v2.2.0 Japanese Canonical Normalization.
Repository:
https://github.com/masahirosakae/industrial-agent-benchmark
Hugging Face Dataset Repository:
https://huggingface.co/datasets/MSakae/industrial-agent-benchmark… See the full description on the dataset page: https://huggingface.co/datasets/MSakae/industrial-agent-benchmark.agent-memory-benchmark
Agent Memory Compression & Evaluation Benchmark
This dataset is a controlled evaluation testbed designed to benchmark long-term memory architectures for conversational AI agents. It stress-tests how agents handle long conversations with complex fact dynamics.
Dataset Structure
1. conversation.json
A 100-turn synthetic conversation (50 user, 50 assistant turns) containing embedded facts categorized under:
Simple Facts: Baseline retrieval details.… See the full description on the dataset page: https://huggingface.co/datasets/kushalicious/agent-memory-benchmark.ai-agent-benchmark
Benchmark Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ai-agent-benchmark.han-cross-agent-skill-transfer-benchmark-dataset-v1
Humanoid Cross-Agent Skill Transfer Benchmark Dataset
This dataset benchmarks how effectively
skills learned by one humanoid agent
can be transferred to another agent
within a decentralized cognitive network.
Objective
To measure cross-agent generalization,
adaptation speed, and transfer efficiency.
Data Fields
source_agent_skill_profile
target_agent_initial_profile
transferred_skill_vector
adaptation_steps
performance_improvement_percentage… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-cross-agent-skill-transfer-benchmark-dataset-v1.Customer_Centric_Agent_Benchmark_C0.1
Comfort. Ease. Joy. Yours.
🆕 Customer-Centric in C0.1
Update
Description
😊 Customer Personalized Comfort
An integrated agent system for comforting customers, delivering 200+ personalized comforts via open-source RL agents—180x cheaper than GPT-5.2-Pro.
🧩 Customer Satisfaction Score
A customer-centric score to measure satisfaction in comfort-focused customers, prioritizing comfort over accuracy performance.
🧩 Benchmarking in Comparison with SOTAs… See the full description on the dataset page: https://huggingface.co/datasets/deepgo/Customer_Centric_Agent_Benchmark_C0.1.
