datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MCP-AtlasMCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
Leaderboard | MCP Atlas Paper | Github
Dataset Summary
This public release is a subset of 500 sample tasks from the MCP Atlas Benchmark dataset.
MCP Atlas is a large-scale benchmark for evaluating tool-use competency, comprising 36 real MCP servers and 220 tools.
Tasks are designed to assess tool-use competency in realistic, multi-step workflows.
Tasks use natural language prompts that avoid… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/MCP-Atlas.MCP-AtlasMCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
Leaderboard | MCP Atlas Paper | Github
Dataset Summary
This public release is a subset of 500 sample tasks from the MCP Atlas Benchmark dataset.
MCP Atlas is a large-scale benchmark for evaluating tool-use competency, comprising 36 real MCP servers and 220 tools.
Tasks are designed to assess tool-use competency in realistic, multi-step workflows.
Tasks use natural language prompts that avoid… See the full description on the dataset page: https://huggingface.co/datasets/txchmechanicus/MCP-Atlas.mcp-clients
MCP Clients Dataset
MCP client identity and capability observations from huggingface.co/mcp.
The pipeline incrementally processes completed daily source partitions. Each
release records its source revision and watermark in state/mcp-clients-v1.json
and a sanitized 60-day dashboard snapshot in reports/dashboard-v1.json.
Its window is the requested UTC date range. Omitted client traffic dates are
unavailable data, not zero traffic.
Protocol traffic before 2026-07-27 is a one-time… See the full description on the dataset page: https://huggingface.co/datasets/evalstate/mcp-clients.mcp_atlasplaywright-mcp-toolcalling
Purpose
I wanted to train a small agent to use a browser effectively, most smaller models I tried <32b struggled to call the tools correctly.
I created this dataset for two main reasons:
To help with finetuning smaller models to use the browser specific tools in playwright.
To look at the security implications of giving browser access to untrusted open-weight models, see blog post.
Versions
I am ironing out the kinks, but I will leave the older versions here in… See the full description on the dataset page: https://huggingface.co/datasets/jdaddyalbs/playwright-mcp-toolcalling.mcptoplist
MCP Ecosystem Dataset
A daily-refreshed, research-friendly snapshot of the Model Context Protocol (MCP)
server ecosystem, published by mcptoplist.com. It covers
every server tracked across the five major MCP registries — the Official MCP Registry,
Glama, Smithery, mcp.so, and PulseMCP — deduplicated to canonical entities, with
per-registry membership and daily registry growth series.
Current version: v2026-09-26 (generated 2026-09-26)
License: CC BY 4.0 — free to use
with… See the full description on the dataset page: https://huggingface.co/datasets/BIFF-AI/mcptoplist.dummy_mcpmcp-ecosystem-sample
MCP Ecosystem - sample of 20 servers and 20 consumers
A preview of a larger dataset on the Model Context Protocol ecosystem: 20 MCP
servers and 20 repositories that use MCP, each with its own tools,
dependencies, registry listings, issues, pull requests, reviews and commits.
6,053 rows across 19 tables, as .parquet and .csv.
Schema is identical to the full dataset. No columns were added.
Inclusion criteria
Keyword search for "mcp" is noisy — a repository can match… See the full description on the dataset page: https://huggingface.co/datasets/mcp-research-alliance/mcp-ecosystem-sample.limbic-eval-tool-use-mcp
Dataset Summary
The MCP Tool Call Evaluation Test Dataset is a synthetic dataset designed for evaluating and benchmarking language models' ability to correctly execute function calls in the context of Model Context Protocol (MCP) tools. This dataset contains 9,813 test examples that assess a model's proficiency in:
Tool Selection: Choosing the correct function from available tools
Parameter Structure: Providing all required parameters with correct names
Parameter Values: Supplying… See the full description on the dataset page: https://huggingface.co/datasets/quotientai/limbic-eval-tool-use-mcp.smoltrace-finance-tasks
SMOLTRACE Synthetic Dataset
This dataset was generated using the TraceMind MCP Server's synthetic data generation tools.
Dataset Info
Tasks: 100
Format: SMOLTRACE evaluation format
Generated: AI-powered synthetic task generation
Usage with SMOLTRACE
from datasets import load_dataset
# Load dataset
dataset = load_dataset("MCP-1st-Birthday/smoltrace-finance-tasks")
# Use with SMOLTRACE
# smoltrace-eval --model openai/gpt-4 --dataset-name… See the full description on the dataset page: https://huggingface.co/datasets/MCP-1st-Birthday/smoltrace-finance-tasks.omnimcp_mcp_prompt_injection_guard_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_prompt_injection_guard_teaser.omnimcp_agentselfheal_mcp_pro_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_agentselfheal_mcp_pro_teaser.omnimcp_mcp_privilege_escalation_auditor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_privilege_escalation_auditor_teaser.MCPWorldsmoltrace-logistics-tasks
SMOLTRACE Synthetic Dataset
This dataset was generated using the TraceMind MCP Server's synthetic data generation tools.
Dataset Info
Tasks: 100
Format: SMOLTRACE evaluation format
Generated: AI-powered synthetic task generation
Usage with SMOLTRACE
from datasets import load_dataset
# Load dataset
dataset = load_dataset("MCP-1st-Birthday/smoltrace-logistics-tasks")
# Use with SMOLTRACE
# smoltrace-eval --model openai/gpt-4 --dataset-name… See the full description on the dataset page: https://huggingface.co/datasets/MCP-1st-Birthday/smoltrace-logistics-tasks.omnimcp_mcp_ssrf_egress_firewall_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_ssrf_egress_firewall_teaser.mcp-fbas
MCP Falsely-Benign Attack (FBA) & Truly-Benign (TB) Preference Dataset
TL;DR
Most LLM safety training targets prompts that look malicious. This dataset targets prompts
that don't. It contains preference pairs for training refusal guardrails against
falsely benign attacks (FBAs) — Model Context Protocol (MCP) tool-use exploits derived
from real CVEs, phrased as ordinary, harmless-sounding requests with no refusal-triggering
language, paired with truly-benign… See the full description on the dataset page: https://huggingface.co/datasets/johnhalloran/mcp-fbas.omnimcp_mcp_filesystem_sandbox_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_filesystem_sandbox_teaser.mcp_toolcall_sandbox_escape_guard_teaser
🚀 AI Safety - MCP Tool-Calling Security & Sandbox Escape Guard (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (332 Samples) & Commercial EULA on Gumroad:👉 AI Safety - MCP Tool-Calling Security & Sandbox Escape Guard on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
📦 What is Inside the Full Production Package:
332 Verified FAANG v2.0 Scenarios (100% AST-Valid Python)… See the full description on the dataset page: https://huggingface.co/datasets/emgena/mcp_toolcall_sandbox_escape_guard_teaser.mcp-tool-use-quality-benchmark
📝 Dataset Summary
The Tool Selection Quality Benchmark dataset evaluates the correctness of function calls generated by an AI assistant 🤖, given:
A set of available tools 🛠️
A conversation history 💬
This dataset is designed to measure Tool Selection Quality.
📂 Dataset Structure
📑 Columns
Column Name
Type
Description
tools_list
string
JSON list of available tools, including names, descriptions, and parameters.
messages_history
string… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/mcp-tool-use-quality-benchmark.mcp-server-bench-gradio-optimized
🔬 Gradio vs FastMCP Benchmark Report
Generated: 2026-03-02T13:04:10.857215
Total scenarios: 48
Executive Summary
echo: Fastmcp wins (96.6 vs 176.5 RPS, 1.83x difference)
Gradio best config: concurrency_limit=nan
fibonacci: Fastmcp wins (43.5 vs 57.1 RPS, 1.31x difference)
Gradio best config: concurrency_limit=nan
async_sleep: Gradio wins (93.1 vs 80.2 RPS, 1.16x difference)
Gradio best config: concurrency_limit=nan
payload_echo: Fastmcp wins (82.5 vs 164.8 RPS… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/mcp-server-bench-gradio-optimized.omnimcp_mcp_github_issue_pr_ops_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_github_issue_pr_ops_teaser.omnimcp_mcp_protocol_handshake_router_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_protocol_handshake_router_teaser.omnimcp_mcp_brave_web_search_triage_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_brave_web_search_triage_teaser.mcp-server-bench
🔬 Gradio vs FastMCP Benchmark Report
Generated: 2026-02-28T05:20:52.935313
Total scenarios: 360
Executive Summary
echo: Fastmcp wins (48.2 vs 192.8 RPS, 4.0x difference)
Gradio best config: concurrency_limit=10.0
fibonacci: Fastmcp wins (36.2 vs 58.3 RPS, 1.61x difference)
Gradio best config: concurrency_limit=5.0
json_transform: Fastmcp wins (50.8 vs 174.8 RPS, 3.44x difference)
Gradio best config: concurrency_limit=nan
async_sleep: Fastmcp wins (59.6 vs 99.2 RPS… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/mcp-server-bench.gradio-agents-mcp-hackathon-certificates
Dataset Card for "gradio-agents-mcp-hackathon-certificates"
More Information needed
mcp-tool-use-evalsmoltrace-recruitment-tasks
SMOLTRACE Synthetic Dataset
This dataset was generated using the TraceMind MCP Server's synthetic data generation tools.
Dataset Info
Tasks: 101
Format: SMOLTRACE evaluation format
Generated: AI-powered synthetic task generation
Usage with SMOLTRACE
from datasets import load_dataset
# Load dataset
dataset = load_dataset("MCP-1st-Birthday/smoltrace-recruitment-tasks")
# Use with SMOLTRACE
# smoltrace-eval --model openai/gpt-4 --dataset-name… See the full description on the dataset page: https://huggingface.co/datasets/MCP-1st-Birthday/smoltrace-recruitment-tasks.transformers-knowledge-graphmcp-memory-auto-trigger-ultimate
