datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sql-create-context
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.McEvalMcEval benchmark data as described in the McEval Paper. Code for the evaluation can be found on Github as McEval.
Qwen3.6-35B-A3B-mcr-stage-b
Qwen3.6-35B-A3B — MCR Stage B Corpus (Distributed Reasoning Localization)
First systematic mechanistic-intervention corpus on a hybrid MoE + GDN + Gated-Attention architecture.
📄 Paper: Loop-Intolerance Profiling: Localizing Distributed Reasoning in a Hybrid MoE Architecture via Nine Convergent Intervention Experiments — submitted to arXiv (2026-04-20, in moderation). Final arXiv ID will be added here once approved.
This dataset contains per-token residual-stream activations at… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/Qwen3.6-35B-A3B-mcr-stage-b.mcp-agent-trajectory-benchmark
MCP Agent Trajectory Benchmark
A benchmark dataset of 49 MCP (Model Context Protocol) agent trajectories (38 single-pass + 11 multi-conv) with complete tool-use traces in the ATIF v1.2 (Agent Trajectory Interchange Format) format. Each agent operates in a distinct business domain with custom tools, realistic user conversations, and full execution traces.
Designed for training and evaluating tool-use / function-calling capabilities of LLMs.
Overview
Item
Details… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark.gspc-mcp
GSPC — conformance bank (MCPBench)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the conformance row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=conformance (family, kind, status and n are on that row, never… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-mcp.mcp-atlas-easy
MCP-Atlas-Easy
An easy, single-tool-call benchmark for pretrained (base) language models, derived from ScaleAI/MCP-Atlas.
MCP-Atlas evaluates instruction-tuned agents on multi-step tool orchestration (3–6 calls per task across 36 real MCP servers). MCP-Atlas-Easy strips that down to the simplest possible form of the same skill: one tool spec, one trivially unambiguous request, one correct tool call, then stop. This makes it usable as a completion-style eval for base models with… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/mcp-atlas-easy.McEval-InstructMcEval-Instruct data as described in the McEval Paper. Code for the evaluation and sft can be found on Github as McEval.
mcp-universe-trajectories
MCP-Universe Agent Trajectories — financial_analysis × DeepSeek V4 Pro
Agent rollout trajectories generated by running every task in the
MCP-Universe
financial_analysis benchmark domain (40 tasks) against DeepSeek V4 Pro
through a slime-compatible custom-generate adapter
(slime_mcp_rollout/).
Each trajectory captures the full multi-turn ReAct/function-call loop:
LLM prompts/responses, every tool call (yfinance + calculator), tool
results, the final answer, and an evaluator-based… See the full description on the dataset page: https://huggingface.co/datasets/Shuibai12138/mcp-universe-trajectories.mc4-pt-cleaned
Description
This is a clenned version of AllenAI mC4 PtBR section. The original dataset can be found here https://huggingface.co/datasets/allenai/c4
Clean procedure
We applied the same clenning procedure as explained here: https://gitlab.com/yhavinga/c4nlpreproc.git
The repository offers two strategies. The first one, found in the main.py file, uses pyspark to create a dataframe that can both clean the text and create a
pseudo mix on the entire dataset. We found this… See the full description on the dataset page: https://huggingface.co/datasets/thegoodfellas/mc4-pt-cleaned.jam-actions-v0
Dataset Card for jam-actions-v0 (public subset)
Version: 0.6.0 — a correction release. It withdraws 58 records whose source arrangements could not be licence-cleared and changes no remaining record. See Version 0.6.0 correction.
Records built: 2026-07-11 (0.5.0 cut; unchanged) Package built: 2026-09-25
DOI: 10.5281/zenodo.22961580 (this version; concept DOI 10.5281/zenodo.22961579). Earlier versions: 0.5.0 10.5281/zenodo.21313954 and 0.4.3 10.5281/zenodo.20279919. Both contain… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v0.jam-actions-v1
jam-actions-v1
Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) ·
Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE ·
Source repo: mcp-tool-shop-org/ai-jam-sessions
The successor to jam-actions-v0.
Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from
what the tools return — and it exists in its current shape because, seven training runs in a row,
the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.jam-actions-acoustic-v0
Dataset Card for jam-actions-acoustic-v0
Version: 1.1.0
Published at mcp-tool-shop/jam-actions-acoustic-v0. No DOI.
Summary
72 constructible gold records of grounded MCP tool use over monophonic audio analysis. Each record pairs a 4-note right-hand reduction of a public-domain library phrase with a seeded synthetic take and a gold verdict (match, pitch fail/warn, timing fail/pass, missed, extra, in-tune vibrato, or nothing-to-grade silence).
This is not a musical… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-acoustic-v0.McKinsey-Reportsmeta-llama/synthetic-data-kit
https://github.com/meta-llama/synthetic-data-kit
McKinsey reports
https://www.mckinsey.com/featured-insights/insights-store
mcp-agent-trajectory-benchmark
⚡ Model Context Protocol (MCP) & Advanced Tool‑Use Alignment Tiers
Official Enterprise Data Repository by springofwindslabs
👉 Looking for full production data?
The complete 1,000-row standard volume and 2,300+ row mutually exclusive, non-overlapping extended package are fully available for commercial deployment via our official procurement gateway:
➔… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/mcp-agent-trajectory-benchmark.bundledex-okf-bundles
BundleDex — OKF Bundle Directory Dataset
A curated directory of 249 Open Knowledge Format (OKF) bundles for AI agents and knowledge systems. Each entry represents a GitHub repository that publishes OKF-formatted knowledge.
Dataset Structure
Each row is a bundle with:
Field
Type
Description
slug
string
Unique identifier
name
string
Human-readable name
description
string
What the bundle covers
author
string
GitHub author/owner
stars
integer
GitHub… See the full description on the dataset page: https://huggingface.co/datasets/Rex-McClawd/bundledex-okf-bundles.cli-commands-explained
Overview
This dataset is a collection of 16,098 command line instructions sourced from Commandlinefu and Cheatsheets. It includes an array of commands, each with an id, title, description, date, url to source, author, votes, and flag indicating if the description is AI generated. The descriptions are primarily authored by the original contributors, for entries where descriptions were absent, they have been generated using NeuralBeagle14-7B. Out of the total entries, 10,039… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/cli-commands-explained.LUNA-RAG-MCP-SFT-10M
Dataset Card for LUNA RAG + MCP SFT Dataset
A clean, English-only, instruction-finetuning dataset for teaching small language models two of the most important 2025–2026 agentic-AI topics: Retrieval-Augmented Generation (RAG) and the Model Context Protocol (MCP).
This repository is the instruction-tuning (SFT) companion dataset for the LUNA-100M model family. It is intentionally compact (≈10M formatted tokens, ≤1,024 tokens per sample) so that it can be absorbed efficiently by a… See the full description on the dataset page: https://huggingface.co/datasets/ASTERIZER/LUNA-RAG-MCP-SFT-10M.jam-rollout-arc-evals
Rollout arc — raw generations
Every model generation behind the write-ups in
mcp-tool-shop-org/ai-jam-sessions
under experiments/rollout-arc/p4/.
Two things you can do with this.
Check our arithmetic. The repo has the readout scripts, the preregistrations and the
intervals — but the generations they were computed from are ~51 MB and were never committed, so
a clone got the conclusions and no way to recompute them. These are those files, unfiltered.
Or run the loop yourself. The… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-rollout-arc-evals.mc2_corpus
MC^2: A Multilingual Corpus of Minority Languages in China
We present MC^2, a Multilingual Corpus of Minority Languages in China, which is the largest open-source corpus so far. This corpus encompasses four languages, namely Tibetan, Uyghur, Kazakh written in the Kazakh Arabic script, and Mongolian written in the traditional Mongolian script.
Please read our paper for more information: MC^2: Towards Transparent and Culturally-Aware NLP for Minority Languages in China (ACL 2024).
The… See the full description on the dataset page: https://huggingface.co/datasets/pkupie/mc2_corpus.vi_grade_school_math_mcq
Dataset Card for Vietnamese Grade School Math Dataset
Dataset Summary
The dataset includes multiple-choice math exercises for elementary school students from grades 1 to 5 in Vietnam.
Supported Tasks and Leaderboards
Languages
The majority of the data is in Vietnamese.
Dataset Structure
Data Instances
The data includes information about the page paths we crawled and some text that has been post-processed. The structure will be… See the full description on the dataset page: https://huggingface.co/datasets/hllj/vi_grade_school_math_mcq.m2mcent-mcp-schemas
🌐 M2MCent Agentic Services - MCP Schemas Dataset
🚀 Empowering Autonomous AI on Base L2
This dataset contains the JSON schemas for 1,005 microservices natively available on the M2MCent Network via the x402 V2 Protocol (EIP-3009).
It is specifically designed for instruction-tuning LLMs (like Llama-3, Mistral, Qwen) so they can autonomously discover, negotiate, and consume monetized API endpoints using gasless cryptocurrency settlements on the Base L2 network.… See the full description on the dataset page: https://huggingface.co/datasets/evozim/m2mcent-mcp-schemas.mc4-zh-idiom-cpt
mC4 zh — Idiom-Tagged Continued-Pretraining Corpus
A 9.6M-document Chinese corpus for continued pretraining on cultural knowledge in
figurative language. Each document is natural web text (from the C4/mC4 zh subset)
containing at least one culturally meaningful chengyu, with an appended knowledge
block that lists every matched idiom together with its figurative meaning(s) and
classical source citation.
Built 2026-07-16 as Stage 1 (continue-pretraining data) of the… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/mc4-zh-idiom-cpt.kubectl-mcp-server-tool-call-reasoning-6k
kubectl-mcp-server-tool-call-reasoning-6k
MCP tool-calling SFT 資料集,由 Agent Tools Fine-Tuning Platform 以「反向生成 + teacher solver 驗證」流程產生。
語言:繁體中文
工具(來自 MCP server):install_helm_chart, upgrade_helm_chart, uninstall_helm_chart, helm_list, helm_status, helm_history, helm_get_values, helm_get_manifest, helm_get_notes, helm_get_hooks, helm_get_all, helm_show_chart, helm_show_values, helm_show_readme, helm_show_crds, helm_show_all, helm_search_repo, helm_search_hub, helm_repo_list… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/kubectl-mcp-server-tool-call-reasoning-6k.stambecco_data_it
🌁 Stambecco-Cleaned: Italian Instruction-Tuning Dataset
The Stambecco-Cleaned Dataset is an Italian translation and adaptation of the community-curated Alpaca-Cleaned dataset, created to enable and evaluate instruction-following capabilities in Italian Large Language Models (LLMs).
📌 Dataset Summary
Language: Italian (it)
Base Source: Alpaca-Cleaned (curated version of Stanford Alpaca)
Primary Use Case: Instruction fine-tuning, evaluation, and alignment for… See the full description on the dataset page: https://huggingface.co/datasets/mchl-labs/stambecco_data_it.mcp-tool-traces-v2
MCP Tool Traces v2 (2K)
Realistic Model Context Protocol (MCP) tool traces with explicit reasoning chains and multi-server orchestration.
What is MCP?
Model Context Protocol is Anthropic's open standard for connecting LLMs to external tools and data sources. MCP servers expose tools that AI assistants can call to interact with filesystems, databases, APIs, and services.
Dataset Description
2,000 traces across 8 real-world MCP servers:
Server… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/mcp-tool-traces-v2.math-for-mcpNursData-MCQ
NursData-MCQ
Dataset Description
NursData-MCQ is a Chinese nursing multiple-choice benchmark for evaluating large language models in fundamental nursing knowledge and nursing-domain reasoning. It was used as the automated evaluation dataset in the EviNurse study.
EviNurse is a domain-specific large language model for evidence-based nursing, developed on Qwen3-32B with supervised fine-tuning and retrieval-augmented generation. In the manuscript, automated… See the full description on the dataset page: https://huggingface.co/datasets/Agnania/NursData-MCQ.kasa-mcp-indirect-channel-probes
KASA MCP — Indirect-Channel Agent Probes
Four probes measuring whether untrusted content — not the operator — can steer a local model that sits inside an agent pipeline. Four model configurations, five runs each, 80 rows.
The dataset exists because it caught a failure in the architecture that produced it. The headline result is A8: 20 out of 20 runs compromised, on every configuration tested.
What each probe measures
Probe
Channel
Question… See the full description on the dataset page: https://huggingface.co/datasets/Earthen937/kasa-mcp-indirect-channel-probes.workix-mcp
Workix — knowledge and tool-use dataset
What Workix is, and how an AI agent talks to it. Package @workix/mcp, hub workix.co,
MCP registry id co.workix/mcp.
Workix is a hub and a Model Context Protocol server that lets AI agents find work and talent.
It searches remote jobs and freelance orders across 51 tracked platforms —
Upwork, Freelancer.com, hh.ru, Kwork, Freelancehunt, RemoteOK, Remotive, Himalayas,
We Work Remotely, Adzuna, Indeed, Glassdoor, ZipRecruiter, Naukri, Jobicy… See the full description on the dataset page: https://huggingface.co/datasets/workix/workix-mcp.flowjudge-dialam
FlowJudge DialAM incremental argument patches
This is the actual transformed dataset used to test whether a 0.6B open model
can learn one falsifiable behavior: given one new proposition and one complete
fixed-size block of earlier propositions from the same dialogue, emit every and
only direct SUPPORT, ATTACK, or REPHRASE edge as one bare JSON object.
An empty relation list is required when no direct edge exists.
The package includes the nested v1 data-efficiency curve, the v2… See the full description on the dataset page: https://huggingface.co/datasets/mr-mc/flowjudge-dialam.
