datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mcp-agent-trajectory-benchmark
MCP Agent Trajectory Benchmark
A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces.
Designed for training and evaluating tool-use / function-calling capabilities of LLMs.
Overview
Item
Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.mcp-atlas-easy
MCP-Atlas-Easy
An easy, single-tool-call benchmark for pretrained (base) language models, derived from ScaleAI/MCP-Atlas.
MCP-Atlas evaluates instruction-tuned agents on multi-step tool orchestration (3–6 calls per task across 36 real MCP servers). MCP-Atlas-Easy strips that down to the simplest possible form of the same skill: one tool spec, one trivially unambiguous request, one correct tool call, then stop. This makes it usable as a completion-style eval for base models with… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/mcp-atlas-easy.gspc-mcp
GSPC — conformance bank (MCPBench)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the conformance row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=conformance (family, kind, status and n are on that row, never… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-mcp.mcp-universe-trajectories
MCP-Universe Agent Trajectories — financial_analysis × DeepSeek V4 Pro
Agent rollout trajectories generated by running every task in the
MCP-Universe
financial_analysis benchmark domain (40 tasks) against DeepSeek V4 Pro
through a slime-compatible custom-generate adapter
(slime_mcp_rollout/).
Each trajectory captures the full multi-turn ReAct/function-call loop:
LLM prompts/responses, every tool call (yfinance + calculator), tool
results, the final answer, and an evaluator-based… See the full description on the dataset page: https://huggingface.co/datasets/Shuibai12138/mcp-universe-trajectories.jam-actions-v1
jam-actions-v1
Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) ·
Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE ·
Source repo: mcp-tool-shop-org/ai-jam-sessions
The successor to jam-actions-v0.
Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from
what the tools return — and it exists in its current shape because, seven training runs in a row,
the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.jam-actions-acoustic-v0
Dataset Card for jam-actions-acoustic-v0
Version: 1.0.2
Published at mcp-tool-shop/jam-actions-acoustic-v0. No DOI.
Summary
108 constructible gold records of grounded MCP tool use over monophonic audio analysis. Each record pairs a 4-note right-hand reduction of a public-domain library phrase with a seeded synthetic take and a gold verdict (match, pitch fail/warn, timing fail/pass, missed, extra, in-tune vibrato, or nothing-to-grade silence).
This is not a musical… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-acoustic-v0.mcp-agent-trajectory-benchmark
⚡ Model Context Protocol (MCP) & Advanced Tool‑Use Alignment Tiers
Official Enterprise Data Repository by springofwindslabs
👉 Looking for full production data?
The complete 1,000-row standard volume and 2,300+ row mutually exclusive, non-overlapping extended package are fully available for commercial deployment via our official procurement gateway:
➔… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/mcp-agent-trajectory-benchmark.jam-actions-v0
Dataset Card for jam-actions-v0 (public subset)
Version: 0.5.1 — a documentation-only revision of the 0.5.0 record cut. No record, split or eval artifact changed; the card gained the fine-tuning evaluation banner and the "What's in a record" walkthrough, which had been added on Hugging Face and lived nowhere else.
Records built: 2026-07-11 Source tag: jam-actions-v0-0.5.0-cut-2026-07-11 (record-content correction release — Bach BWV 846 errata 001 + 002; see RELEASE_NOTES.md… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v0.LUNA-RAG-MCP-SFT-10M
Dataset Card for LUNA RAG + MCP SFT Dataset
A clean, English-only, instruction-finetuning dataset for teaching small language models two of the most important 2025–2026 agentic-AI topics: Retrieval-Augmented Generation (RAG) and the Model Context Protocol (MCP).
This repository is the instruction-tuning (SFT) companion dataset for the LUNA-100M model family. It is intentionally compact (≈10M formatted tokens, ≤1,024 tokens per sample) so that it can be absorbed efficiently by a… See the full description on the dataset page: https://huggingface.co/datasets/ASTERIZER/LUNA-RAG-MCP-SFT-10M.jam-rollout-arc-evals
Rollout arc — raw generations
Every model generation behind the write-ups in
mcp-tool-shop-org/ai-jam-sessions
under experiments/rollout-arc/p4/.
Two things you can do with this.
Check our arithmetic. The repo has the readout scripts, the preregistrations and the
intervals — but the generations they were computed from are ~51 MB and were never committed, so
a clone got the conclusions and no way to recompute them. These are those files, unfiltered.
Or run the loop yourself. The… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-rollout-arc-evals.omnimcp_mcp_prompt_injection_guard_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_prompt_injection_guard_teaser.omnimcp_agentselfheal_mcp_pro_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_agentselfheal_mcp_pro_teaser.m2mcent-mcp-schemas
🌐 M2MCent Agentic Services - MCP Schemas Dataset
🚀 Empowering Autonomous AI on Base L2
This dataset contains the JSON schemas for 1,005 microservices natively available on the M2MCent Network via the x402 V2 Protocol (EIP-3009).
It is specifically designed for instruction-tuning LLMs (like Llama-3, Mistral, Qwen) so they can autonomously discover, negotiate, and consume monetized API endpoints using gasless cryptocurrency settlements on the Base L2 network.… See the full description on the dataset page: https://huggingface.co/datasets/evozim/m2mcent-mcp-schemas.omnimcp_mcp_privilege_escalation_auditor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_privilege_escalation_auditor_teaser.perception-mcp-benchmark
Perception MCP Workflow Benchmark
A side-by-side benchmark of AI assistants doing real digital-asset research workflows, with and without the Perception MCP connected.
Question: does connecting an industry-specific data corpus to a frontier AI assistant produce measurably better research work than the same assistant with its native web search?
Answer, across 48 scored runs: yes. Blind-judged mean score 11.4 → 16.4 (max 25, +44%), with the widest gains in recency (+74%) and… See the full description on the dataset page: https://huggingface.co/datasets/ferniko/perception-mcp-benchmark.omnimcp_mcp_ssrf_egress_firewall_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_ssrf_egress_firewall_teaser.mcp-fbas
MCP Falsely-Benign Attack (FBA) & Truly-Benign (TB) Preference Dataset
TL;DR
Most LLM safety training targets prompts that look malicious. This dataset targets prompts
that don't. It contains preference pairs for training refusal guardrails against
falsely benign attacks (FBAs) — Model Context Protocol (MCP) tool-use exploits derived
from real CVEs, phrased as ordinary, harmless-sounding requests with no refusal-triggering
language, paired with truly-benign… See the full description on the dataset page: https://huggingface.co/datasets/johnhalloran/mcp-fbas.omnimcp_mcp_filesystem_sandbox_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_filesystem_sandbox_teaser.omnimcp_mcp_github_issue_pr_ops_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_github_issue_pr_ops_teaser.omnimcp_mcp_protocol_handshake_router_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_protocol_handshake_router_teaser.omnimcp_mcp_brave_web_search_triage_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_brave_web_search_triage_teaser.mcp_toolcall_sandbox_escape_guard_teaser
🚀 AI Safety - MCP Tool-Calling Security & Sandbox Escape Guard (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (332 Samples) & Commercial EULA on Gumroad:👉 AI Safety - MCP Tool-Calling Security & Sandbox Escape Guard on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
📦 What is Inside the Full Production Package:
332 Verified FAANG v2.0 Scenarios (100% AST-Valid Python)… See the full description on the dataset page: https://huggingface.co/datasets/emgena/mcp_toolcall_sandbox_escape_guard_teaser.kubectl-mcp-server-tool-call-reasoning-6k
kubectl-mcp-server-tool-call-reasoning-6k
MCP tool-calling SFT 資料集,由 Agent Tools Fine-Tuning Platform 以「反向生成 + teacher solver 驗證」流程產生。
語言:繁體中文
工具(來自 MCP server):install_helm_chart, upgrade_helm_chart, uninstall_helm_chart, helm_list, helm_status, helm_history, helm_get_values, helm_get_manifest, helm_get_notes, helm_get_hooks, helm_get_all, helm_show_chart, helm_show_values, helm_show_readme, helm_show_crds, helm_show_all, helm_search_repo, helm_search_hub, helm_repo_list… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/kubectl-mcp-server-tool-call-reasoning-6k.mcp-tool-traces-v2
MCP Tool Traces v2 (2K)
Realistic Model Context Protocol (MCP) tool traces with explicit reasoning chains and multi-server orchestration.
What is MCP?
Model Context Protocol is Anthropic's open standard for connecting LLMs to external tools and data sources. MCP servers expose tools that AI assistants can call to interact with filesystems, databases, APIs, and services.
Dataset Description
2,000 traces across 8 real-world MCP servers:
Server… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/mcp-tool-traces-v2.math-for-mcpworkix-mcp
Workix — knowledge and tool-use dataset
What Workix is, and how an AI agent talks to it. Package @workix/mcp, hub workix.co,
MCP registry id co.workix/mcp.
Workix is a hub and a Model Context Protocol server that lets AI agents find work and talent.
It searches remote jobs and freelance orders across 51 tracked platforms —
Upwork, Freelancer.com, hh.ru, Kwork, Freelancehunt, RemoteOK, Remotive, Himalayas,
We Work Remotely, Adzuna, Indeed, Glassdoor, ZipRecruiter, Naukri, Jobicy… See the full description on the dataset page: https://huggingface.co/datasets/workix/workix-mcp.kasa-mcp-indirect-channel-probes
KASA MCP — Indirect-Channel Agent Probes
Four probes measuring whether untrusted content — not the operator — can steer a local model that sits inside an agent pipeline. Four model configurations, five runs each, 80 rows.
The dataset exists because it caught a failure in the architecture that produced it. The headline result is A8: 20 out of 20 runs compromised, on every configuration tested.
What each probe measures
Probe
Channel
Question… See the full description on the dataset page: https://huggingface.co/datasets/Earthen937/kasa-mcp-indirect-channel-probes.mcp-servers-catalog
MCP Servers Catalog
A structured, machine-readable catalog of 2,223 MCP (Model Context Protocol) servers extracted from the curated awesome-mcp-servers list. Covers 49 categories with language, scope, OS, and officiality metadata for each server.
Updated monthly. Last updated: 2026-05.
Dataset Summary
This dataset catalogs what each MCP server is: its name, repository URL, category, description, implementation language, deployment scope (cloud/local), supported… See the full description on the dataset page: https://huggingface.co/datasets/automatelab/mcp-servers-catalog.mcp-tool-traces
mcp-tool-traces
135 synthetic MCP (Model Context Protocol) tool-use traces covering three real SaaS servers: Stripe, Linear, and Notion. Apache 2.0 — commercial use permitted.
The first open dataset of multi-turn agent conversations using actual MCP server schemas. Generated as the MCP specification reaches Release Candidate (July 2026).
Quick Load
from datasets import load_dataset
ds = load_dataset("stindardlogic/mcp-tool-traces", split="train")
# Filter by MCP… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/mcp-tool-traces.kubectl-mcp-server-tool-call-reasoning-1k
Kubernetes MCP Function-Calling SFT Dataset
本資料集是一個專為 Kubernetes 操作設計的高品質 Function-Calling 微調資料集,包含 1,500 筆指令與工具呼叫對應樣本,旨在提升大語言模型在 Kubernetes 集群管理與診斷任務中的自動化執行能力。
資料來源與生成方式
本資料集基於 MCP (Model Context Protocol) 定義的 Kubernetes 工具集生成。我們使用 google/gemini-3.1-flash-lite 作為生成模型,針對各類 Kubernetes 運維場景(如網路策略診斷、資源備份、Pod 健康檢查等)進行合成數據建構。
資料集格式嚴格對齊 twinkle-ai/tw-function-call-reasoning-10k 標準,確保模型在執行工具前具備良好的思維鏈(Chain-of-Thought)推理能力。
品質控管與驗證流程… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/kubectl-mcp-server-tool-call-reasoning-1k.
