datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mcp-agent-trajectory-benchmark
MCP Agent Trajectory Benchmark
A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces.
Designed for training and evaluating tool-use / function-calling capabilities of LLMs.
Overview
Item
Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.mcp-atlas-easy
MCP-Atlas-Easy
An easy, single-tool-call benchmark for pretrained (base) language models, derived from ScaleAI/MCP-Atlas.
MCP-Atlas evaluates instruction-tuned agents on multi-step tool orchestration (3–6 calls per task across 36 real MCP servers). MCP-Atlas-Easy strips that down to the simplest possible form of the same skill: one tool spec, one trivially unambiguous request, one correct tool call, then stop. This makes it usable as a completion-style eval for base models with… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/mcp-atlas-easy.gspc-mcp
GSPC — conformance bank (MCPBench)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the conformance row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=conformance (family, kind, status and n are on that row, never… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-mcp.mcp-universe-trajectories
MCP-Universe Agent Trajectories — financial_analysis × DeepSeek V4 Pro
Agent rollout trajectories generated by running every task in the
MCP-Universe
financial_analysis benchmark domain (40 tasks) against DeepSeek V4 Pro
through a slime-compatible custom-generate adapter
(slime_mcp_rollout/).
Each trajectory captures the full multi-turn ReAct/function-call loop:
LLM prompts/responses, every tool call (yfinance + calculator), tool
results, the final answer, and an evaluator-based… See the full description on the dataset page: https://huggingface.co/datasets/Shuibai12138/mcp-universe-trajectories.jam-actions-v1
jam-actions-v1
Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) ·
Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE ·
Source repo: mcp-tool-shop-org/ai-jam-sessions
The successor to jam-actions-v0.
Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from
what the tools return — and it exists in its current shape because, seven training runs in a row,
the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.jam-actions-acoustic-v0
Dataset Card for jam-actions-acoustic-v0
Version: 1.0.2
Published at mcp-tool-shop/jam-actions-acoustic-v0. No DOI.
Summary
108 constructible gold records of grounded MCP tool use over monophonic audio analysis. Each record pairs a 4-note right-hand reduction of a public-domain library phrase with a seeded synthetic take and a gold verdict (match, pitch fail/warn, timing fail/pass, missed, extra, in-tune vibrato, or nothing-to-grade silence).
This is not a musical… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-acoustic-v0.mcp-agent-trajectory-benchmark
⚡ Model Context Protocol (MCP) & Advanced Tool‑Use Alignment Tiers
Official Enterprise Data Repository by springofwindslabs
👉 Looking for full production data?
The complete 1,000-row standard volume and 2,300+ row mutually exclusive, non-overlapping extended package are fully available for commercial deployment via our official procurement gateway:
➔… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/mcp-agent-trajectory-benchmark.jam-actions-v0
Dataset Card for jam-actions-v0 (public subset)
Version: 0.5.1 — a documentation-only revision of the 0.5.0 record cut. No record, split or eval artifact changed; the card gained the fine-tuning evaluation banner and the "What's in a record" walkthrough, which had been added on Hugging Face and lived nowhere else.
Records built: 2026-07-11 Source tag: jam-actions-v0-0.5.0-cut-2026-07-11 (record-content correction release — Bach BWV 846 errata 001 + 002; see RELEASE_NOTES.md… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v0.LUNA-RAG-MCP-SFT-10M
Dataset Card for LUNA RAG + MCP SFT Dataset
A clean, English-only, instruction-finetuning dataset for teaching small language models two of the most important 2025–2026 agentic-AI topics: Retrieval-Augmented Generation (RAG) and the Model Context Protocol (MCP).
This repository is the instruction-tuning (SFT) companion dataset for the LUNA-100M model family. It is intentionally compact (≈10M formatted tokens, ≤1,024 tokens per sample) so that it can be absorbed efficiently by a… See the full description on the dataset page: https://huggingface.co/datasets/ASTERIZER/LUNA-RAG-MCP-SFT-10M.jam-rollout-arc-evals
Rollout arc — raw generations
Every model generation behind the write-ups in
mcp-tool-shop-org/ai-jam-sessions
under experiments/rollout-arc/p4/.
Two things you can do with this.
Check our arithmetic. The repo has the readout scripts, the preregistrations and the
intervals — but the generations they were computed from are ~51 MB and were never committed, so
a clone got the conclusions and no way to recompute them. These are those files, unfiltered.
Or run the loop yourself. The… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-rollout-arc-evals.m2mcent-mcp-schemas
🌐 M2MCent Agentic Services - MCP Schemas Dataset
🚀 Empowering Autonomous AI on Base L2
This dataset contains the JSON schemas for 1,005 microservices natively available on the M2MCent Network via the x402 V2 Protocol (EIP-3009).
It is specifically designed for instruction-tuning LLMs (like Llama-3, Mistral, Qwen) so they can autonomously discover, negotiate, and consume monetized API endpoints using gasless cryptocurrency settlements on the Base L2 network.… See the full description on the dataset page: https://huggingface.co/datasets/evozim/m2mcent-mcp-schemas.kubectl-mcp-server-tool-call-reasoning-6k
kubectl-mcp-server-tool-call-reasoning-6k
MCP tool-calling SFT 資料集,由 Agent Tools Fine-Tuning Platform 以「反向生成 + teacher solver 驗證」流程產生。
語言:繁體中文
工具(來自 MCP server):install_helm_chart, upgrade_helm_chart, uninstall_helm_chart, helm_list, helm_status, helm_history, helm_get_values, helm_get_manifest, helm_get_notes, helm_get_hooks, helm_get_all, helm_show_chart, helm_show_values, helm_show_readme, helm_show_crds, helm_show_all, helm_search_repo, helm_search_hub, helm_repo_list… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/kubectl-mcp-server-tool-call-reasoning-6k.mcp-tool-traces-v2
MCP Tool Traces v2 (2K)
Realistic Model Context Protocol (MCP) tool traces with explicit reasoning chains and multi-server orchestration.
What is MCP?
Model Context Protocol is Anthropic's open standard for connecting LLMs to external tools and data sources. MCP servers expose tools that AI assistants can call to interact with filesystems, databases, APIs, and services.
Dataset Description
2,000 traces across 8 real-world MCP servers:
Server… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/mcp-tool-traces-v2.math-for-mcpworkix-mcp
Workix — knowledge and tool-use dataset
What Workix is, and how an AI agent talks to it. Package @workix/mcp, hub workix.co,
MCP registry id co.workix/mcp.
Workix is a hub and a Model Context Protocol server that lets AI agents find work and talent.
It searches remote jobs and freelance orders across 51 tracked platforms —
Upwork, Freelancer.com, hh.ru, Kwork, Freelancehunt, RemoteOK, Remotive, Himalayas,
We Work Remotely, Adzuna, Indeed, Glassdoor, ZipRecruiter, Naukri, Jobicy… See the full description on the dataset page: https://huggingface.co/datasets/workix/workix-mcp.kasa-mcp-indirect-channel-probes
KASA MCP — Indirect-Channel Agent Probes
Four probes measuring whether untrusted content — not the operator — can steer a local model that sits inside an agent pipeline. Four model configurations, five runs each, 80 rows.
The dataset exists because it caught a failure in the architecture that produced it. The headline result is A8: 20 out of 20 runs compromised, on every configuration tested.
What each probe measures
Probe
Channel
Question… See the full description on the dataset page: https://huggingface.co/datasets/Earthen937/kasa-mcp-indirect-channel-probes.mcp-tool-traces
mcp-tool-traces
135 synthetic MCP (Model Context Protocol) tool-use traces covering three real SaaS servers: Stripe, Linear, and Notion. Apache 2.0 — commercial use permitted.
The first open dataset of multi-turn agent conversations using actual MCP server schemas. Generated as the MCP specification reaches Release Candidate (July 2026).
Quick Load
from datasets import load_dataset
ds = load_dataset("stindardlogic/mcp-tool-traces", split="train")
# Filter by MCP… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/mcp-tool-traces.AINeed_mcp
