agent-benchmark
mcp-agent-trajectory-benchmark
MCP Agent Trajectory Benchmark
A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces.
Designed for training and evaluating tool-use / function-calling capabilities of LLMs.
Overview
Item
Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.ii-agent_gaia-benchmark_validationb3-agent-security-benchmark-weak[paper] [blogpost] [game]
b3 AI Security Benchmark: Breaking Agent Backbones
Highly contextalized prompt injections crowd-sourced during the Gandalf Agent Breaker Challenge.
This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security
of Backbone LLMs in AI Agents.
The high quality dataset was used to evaluate the security of more than 30 LLMs.
Dataset Summary
Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.agent-benchmark
configs:
- config_name: default
data_files:
- path: run-log.json
split: train
language:
- en
license: cc-by-nc-4.0
tags:
- ai-agents
- benchmark
- evaluation
- agent-evaluation
- llm-agents
- trustworthy-ai
- product-review
- c2pa
pretty_name: Hlido AI Agent Benchmark
size_categories:
- n<1K
task_categories:
- other
Hlido AI Agent Benchmark
Independent, cryptographically-attested evaluations of AI agents and products.
Hlido reviews AI… See the full description on the dataset page: https://huggingface.co/datasets/hlido-eu/agent-benchmark.finance_agent_benchmark
Finance Agent Benchmark Dataset
We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings.
We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/vals-ai/finance_agent_benchmark.structured-file-audit-benchmark
Paper Data Release
This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them.
Contents
datasets/
Benchmark data and per-task manifests for the three paper-facing splits.
datasets/sc_flat/data
SC-Flat is derived from DaBench, augmented with a replayable perturbation
injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.
