datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
test-mcp-logs(Put queries first as heuristics don't detect when there are no logs)
MCP-AtlasMCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
Leaderboard | MCP Atlas Paper | Github
Dataset Summary
This public release is a subset of 500 sample tasks from the MCP Atlas Benchmark dataset.
MCP Atlas is a large-scale benchmark for evaluating tool-use competency, comprising 36 real MCP servers and 220 tools.
Tasks are designed to assess tool-use competency in realistic, multi-step workflows.
Tasks use natural language prompts that avoid… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/MCP-Atlas.discover-toolsmcpmark-trajectory-log
MCPMark Trajectory Logs (mcpmark-v1-0905)
These logs are publicly available for research use. They capture end-to-end trajectories (prompts, tool calls, and outcomes) for MCPMark tasks across multiple MCP services and models.
Contents
Each task run produces a trajectory folder containing three files:
meta.json: Metadata (task id, model, timestamps, status, etc.)
messages.json: Turn-by-turn exchanges including model thoughts/tool calls
execution.log: Key execution-time… See the full description on the dataset page: https://huggingface.co/datasets/Jakumetsu/mcpmark-trajectory-log.bio-mcp-data
Bio-MCP-Data
A repository containing biological datasets that will be used by BIO-MCP MCP (Model Context Protocol) standard.
About
This repository hosts biological data assets formatted to be compatible with the Model Context Protocol, enabling AI models to efficiently access and process biological information. The data is managed using Git Large File Storage (LFS) to handle large biological datasets.
Purpose
Provide standardized biological datasets for AI… See the full description on the dataset page: https://huggingface.co/datasets/longevity-genie/bio-mcp-data.japan-travel-mcp-data
Japan Travel MCP — Data
The runtime data for the japan-travel-mcp
Model Context Protocol server. Comprehensive Japanese travel data for AI agents,
built from public official sources, covering all 47 prefectures and 1,938 local
government entities.
Code lives on GitHub: github.com/ookami0210/japan-travel-mcp
Data lives here. The npm package downloads this dataset on first run.
Why this dataset exists
Japan's tourism information — created to reach the world — is… See the full description on the dataset page: https://huggingface.co/datasets/open-travel/japan-travel-mcp-data.MCP-AtlasMCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
Leaderboard | MCP Atlas Paper | Github
Dataset Summary
This public release is a subset of 500 sample tasks from the MCP Atlas Benchmark dataset.
MCP Atlas is a large-scale benchmark for evaluating tool-use competency, comprising 36 real MCP servers and 220 tools.
Tasks are designed to assess tool-use competency in realistic, multi-step workflows.
Tasks use natural language prompts that avoid… See the full description on the dataset page: https://huggingface.co/datasets/txchmechanicus/MCP-Atlas.mcp-universe-finance-sft
MCP-Universe Finance SFT
Generated dataset (tasks backup + reward=1 trajectories).
mcp-agent-trajectory-benchmark
MCP Agent Trajectory Benchmark
A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces.
Designed for training and evaluating tool-use / function-calling capabilities of LLMs.
Overview
Item
Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.mcp-clients
MCP Clients Dataset
MCP client identity and capability observations from huggingface.co/mcp.
The pipeline incrementally processes completed daily source partitions. Each
release records its source revision and watermark in state/mcp-clients-v1.json
and a sanitized 60-day dashboard snapshot in reports/dashboard-v1.json.
Its window is the requested UTC date range. Omitted client traffic dates are
unavailable data, not zero traffic.
Protocol traffic before 2026-07-27 is a one-time… See the full description on the dataset page: https://huggingface.co/datasets/evalstate/mcp-clients.mcp-registry
Vinkius Connector Registry — Open Data Initiative
Welcome to the Vinkius Open Data Initiative. We are opening access to the Vinkius connector catalog. This repository provides automatically updated documentation for 9,480 unique connectors for AI agents.
Research & Training Applications
This highly structured corpus is designed specifically for AI researchers, data scientists, and language model developers. It provides a robust foundation for advancing artificial… See the full description on the dataset page: https://huggingface.co/datasets/Vinkius/mcp-registry.gspc-mcp
GSPC — conformance bank (MCPBench)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
The live board is the authority
GET https://councilof.ai/api/gspc — quote totals.public_count. This Hub card is a printer of that GET, never a second
engine. If the fetch… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-mcp.playwright-mcp-toolcalling
Purpose
I wanted to train a small agent to use a browser effectively, most smaller models I tried <32b struggled to call the tools correctly.
I created this dataset for two main reasons:
To help with finetuning smaller models to use the browser specific tools in playwright.
To look at the security implications of giving browser access to untrusted open-weight models, see blog post.
Versions
I am ironing out the kinks, but I will leave the older versions here in… See the full description on the dataset page: https://huggingface.co/datasets/jdaddyalbs/playwright-mcp-toolcalling.mcp-atlas-easy
MCP-Atlas-Easy
An easy, single-tool-call benchmark for pretrained (base) language models, derived from ScaleAI/MCP-Atlas.
MCP-Atlas evaluates instruction-tuned agents on multi-step tool orchestration (3–6 calls per task across 36 real MCP servers). MCP-Atlas-Easy strips that down to the simplest possible form of the same skill: one tool spec, one trivially unambiguous request, one correct tool call, then stop. This makes it usable as a completion-style eval for base models with… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/mcp-atlas-easy.rl_rag_sqa_searcharena_rubrics_web_augmented_outcome_with_new_mcp_system_promptmcp_atlasmcpc-modskills_go_to_githubworking here https://github.com/evalstate/skills-dev
mods-mcpevideo-mcp
Video-MCP
Video-MCP is a synthetic video dataset for training and evaluating video generation models on multiple-choice question-answering (MCQA) tasks. Each sample is a short video clip (~5 seconds) where a visual question-answering prompt is embedded directly into the video frames, and the correct answer is revealed by progressively highlighting one of four answer boxes (A/B/C/D) over the duration of the clip.
The… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/video-mcp.mcp-universe-trajectories
MCP-Universe Agent Trajectories — financial_analysis × DeepSeek V4 Pro
Agent rollout trajectories generated by running every task in the
MCP-Universe
financial_analysis benchmark domain (40 tasks) against DeepSeek V4 Pro
through a slime-compatible custom-generate adapter
(slime_mcp_rollout/).
Each trajectory captures the full multi-turn ReAct/function-call loop:
LLM prompts/responses, every tool call (yfinance + calculator), tool
results, the final answer, and an evaluator-based… See the full description on the dataset page: https://huggingface.co/datasets/Shuibai12138/mcp-universe-trajectories.mcptoplist
MCP Ecosystem Dataset
A daily-refreshed, research-friendly snapshot of the Model Context Protocol (MCP)
server ecosystem, published by mcptoplist.com. It covers
every server tracked across the five major MCP registries — the Official MCP Registry,
Glama, Smithery, mcp.so, and PulseMCP — deduplicated to canonical entities, with
per-registry membership and daily registry growth series.
Current version: v2026-09-21 (generated 2026-09-21)
License: CC BY 4.0 — free to use
with… See the full description on the dataset page: https://huggingface.co/datasets/BIFF-AI/mcptoplist.mcphunt-agent-traces
MCPHunt Agent Traces
Agent execution traces from the MCPHunt evaluation framework, measuring
cross-boundary data propagation in multi-server MCP agents.
Contents
main/ — 3,615 traces from 5 models across 147 tasks and 7 environment
variants (risky_v1/v2/v3, benign, hard_neg_v1/v2/v3). One JSON file per model.
mitigation/ — 2,706 traces from the prompt-mitigation study (M0--M3
levels) across 3 models.
live_guard_defense/ — 387 DeepSeek-V4-Flash traces from the… See the full description on the dataset page: https://huggingface.co/datasets/lihaonan0716/mcphunt-agent-traces.mcpmark_winstonMCP-Flowhttps://arxiv.org/abs/2510.24284
mcp-ecosystem-sample
MCP Ecosystem - sample of 20 servers and 20 consumers
A preview of a larger dataset on the Model Context Protocol ecosystem: 20 MCP
servers and 20 repositories that use MCP, each with its own tools,
dependencies, registry listings, issues, pull requests, reviews and commits.
6,053 rows across 19 tables, as .parquet and .csv.
Schema is identical to the full dataset. No columns were added.
Inclusion criteria
Keyword search for "mcp" is noisy — a repository can match… See the full description on the dataset page: https://huggingface.co/datasets/mcp-research-alliance/mcp-ecosystem-sample.rl_rag_sqa_searcharena_rubrics_web_augmented_rubrics_only_with_new_mcp_system_promptdeepfabric-figma-mcp
deepfabric-figma-mcp
Dataset generated with DeepFabric.
sre-8h-mcp-harbor-tasks-v2dummy_mcp
