datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mcp-agent-trajectory-benchmark
MCP Agent Trajectory Benchmark
A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces.
Designed for training and evaluating tool-use / function-calling capabilities of LLMs.
Overview
Item
Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.b3-agent-security-benchmark-weak[paper] [blogpost] [game]
b3 AI Security Benchmark: Breaking Agent Backbones
Highly contextalized prompt injections crowd-sourced during the Gandalf Agent Breaker Challenge.
This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security
of Backbone LLMs in AI Agents.
The high quality dataset was used to evaluate the security of more than 30 LLMs.
Dataset Summary
Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.data-agent-benchmarks
LongHorizon Full Data-Agent Benchmarks
Companion data artifacts for five complete evaluation tracks:
DataSciBench full55 / 167 metric entries
DABStep full450
DABStep-Research full100
DSBench Modeling full74
LongDS full68 / 2,225 turns
The companion GitHub repository contains processed manifests, evaluation code,
historical API ReAct baseline code, download/preparation tools, and the frozen
source lock. artifact_manifest.json records every uploaded object's size,
SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.claude-agent-skills-benchmark
Claude Agent Skills Benchmark
Claude Agent Skills 评测数据集
Description
A benchmark dataset for evaluating whether LLMs can accurately trigger and execute domain-specific Skills on the Claude Code platform. Skills are designed by vertical domain experts with varying complexity levels (based on attachments: scripts, references, assets, and reference markdown files).
Evaluation Scenarios Cover:
Office automation, coding, investment promotion, financial services, industrial… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/claude-agent-skills-benchmark.calendar-agent-benchmark
Calendar Agent Training and Evaluation
Tool Calling SFT와 환경 기반 RLVR을 함께 실습하는 Calendar Agent 데이터입니다.
1. 데이터 구성
파일
개수
구성
SHA-256
sft_train.jsonl
3,000
기본 시나리오, 다중 참석자 생성, 안전 대조군
f882fba3a6c32982d2f8cca06465fa1ec3c68562c655d2ee1bcd92decd29e637
sft_validation.jsonl
300
SFT 학습 중 검증
b657e4162ef26bd580ad975b4f2af575a616dda17c90c7b4edab38790ad66ac5
rlvr_train.jsonl
256
충돌 복구와 다중 참석자 생성 128개, 안전 경계 128개… See the full description on the dataset page: https://huggingface.co/datasets/NotoriousH2/calendar-agent-benchmark.mcp-agent-trajectory-benchmark
⚡ Model Context Protocol (MCP) & Advanced Tool‑Use Alignment Tiers
15-second demo: strict JSONL trajectories + 7-point rubric validation (schema stability 100%).
Schema Validation Summary
Programmatic validation of this exact trial file - reproducible from data.jsonl.
What this trial verifies — use these 50 rows to confirm, on your own stack:
Schema integrity (strict JSONL, matches the published schema)
Multi-turn / tool-use structural consistency… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/mcp-agent-trajectory-benchmark.agent-evaluation-benchmark
Agent Evaluation Benchmark
A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories.
Overview
This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more.
Categories
Category
Test Cases
Description
Weather
5
Forecasts, UV index, climate history
Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.sage-agent-benchmark
SAGE Agent Benchmark
Comprehensive benchmark for evaluating AI agent capabilities across three core competencies:
Tool Selection - Choosing appropriate tools for tasks
Task Planning - Decomposing complex tasks into step sequences
Timing Judgment - Deciding when to use tools vs. direct answers
Dataset Statistics
Total Samples: ~11,000
Tool Selection: ~6,000 samples
Task Planning: ~3,000 samples
Timing Judgment: ~2,000 samples
Splits: train, dev, test
Usage… See the full description on the dataset page: https://huggingface.co/datasets/intellistream/sage-agent-benchmark.agent-sandbox-negotiation-benchmark
Agent Sandbox Negotiation Benchmark v1
Overview
A dataset of simulated multi-agent negotiations generated using the open-source Agent Sandbox framework.
This dataset captures the final negotiation outcomes, turn depths, strategy alignments, and agreed prices of local LLMs (Llama-3 and Mistral) engaged in intense, adversarial price negotiations at massive scale.
Dataset Statistics
Simulations: 24,122
Strategies: 4 (Balanced, Aggressive, Conservative, Adaptive)… See the full description on the dataset page: https://huggingface.co/datasets/ScareRezume/agent-sandbox-negotiation-benchmark.
