datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gspc-mcp
GSPC — conformance bank (MCPBench)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the conformance row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=conformance (family, kind, status and n are on that row, never… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-mcp.nemotron-mc-en-ar-midtrain
nemotron-mc-en-ar-midtrain
Arabic translation of the Nemotron-Pretraining-Multiple-Choice config of Nemotron-Pretraining-Specialized-v1.2 (pinned revision 807afc1). Translated with google/gemma-4-12B-it (bf16, greedy) on A100s. All 23,926,492 source rows are present, none dropped. English source and Arabic translation sit in the same row, so the dataset serves as a parallel corpus as well as an Arabic one. A sibling corpus from the same pipeline is available at… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-mc-en-ar-midtrain.LG_ConvFin_MCQ
LG ConvFinQA MCQ Dataset
Dataset Description
This dataset contains high-quality Multiple-Choice Questions (MCQs) generated from the ConvFinQA dataset for Reward Model (RM) training.
Each question has been:
Generated with 4 carefully crafted answer choices (correct/incorrect × detailed reasoning/answer only)
Verified using hybrid consistency checking:
Self-consistency: N=10 samplings with the same model
Multi-model verification: Cross-validation with 3 different models… See the full description on the dataset page: https://huggingface.co/datasets/ssunggun2/LG_ConvFin_MCQ.mcp-universe-trajectories
MCP-Universe Agent Trajectories — financial_analysis × DeepSeek V4 Pro
Agent rollout trajectories generated by running every task in the
MCP-Universe
financial_analysis benchmark domain (40 tasks) against DeepSeek V4 Pro
through a slime-compatible custom-generate adapter
(slime_mcp_rollout/).
Each trajectory captures the full multi-turn ReAct/function-call loop:
LLM prompts/responses, every tool call (yfinance + calculator), tool
results, the final answer, and an evaluator-based… See the full description on the dataset page: https://huggingface.co/datasets/Shuibai12138/mcp-universe-trajectories.TopiOCQATopiOCQA is an information-seeking conversational dataset with challenging topic switching phenomena.OHB-20k
OHB-20k — Open Human Benchmark
An open-source, community-driven benchmark measuring broad human-facing competence in language
models: not just exam trivia, but the practical, social, civic, and technical reasoning people
actually rely on — plus specialist tracks for long-context recall, hallucination detection, and
grounded table QA.
20,000 questions / 35,000 points / 21 sections (text core)
9 nested size splits (500 → 20,000) so small and large models alike get a fair, cheap… See the full description on the dataset page: https://huggingface.co/datasets/MC7ever/OHB-20k.russian-nmo-medical-mcq
Russian NMO Medical MCQ
Choose language / Выберите язык: Русский | English
Русский
Это датасет русскоязычных медицинских тестовых вопросов НМО с вариантами ответа.
В нем есть вопросы с одним правильным вариантом и вопросы с несколькими правильными
вариантами. Датасет подготовлен так, чтобы его можно было сразу использовать для
тонкой настройки LLM, проверки качества ответов и экспериментов с медицинским QA.
Главная идея простая: дать модели вопрос, тему и варианты ответа… See the full description on the dataset page: https://huggingface.co/datasets/drkolesnikov/russian-nmo-medical-mcq.cli-commands-explained
Overview
This dataset is a collection of 16,098 command line instructions sourced from Commandlinefu and Cheatsheets. It includes an array of commands, each with an id, title, description, date, url to source, author, votes, and flag indicating if the description is AI generated. The descriptions are primarily authored by the original contributors, for entries where descriptions were absent, they have been generated using NeuralBeagle14-7B. Out of the total entries, 10,039… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/cli-commands-explained.jam-rollout-arc-evals
Rollout arc — raw generations
Every model generation behind the write-ups in
mcp-tool-shop-org/ai-jam-sessions
under experiments/rollout-arc/p4/.
Two things you can do with this.
Check our arithmetic. The repo has the readout scripts, the preregistrations and the
intervals — but the generations they were computed from are ~51 MB and were never committed, so
a clone got the conclusions and no way to recompute them. These are those files, unfiltered.
Or run the loop yourself. The… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-rollout-arc-evals.MCiteBench
MCiteBench Dataset
MCiteBench is a benchmark for evaluating the ability of Multimodal Large Language Models (MLLMs) to generate text with citations in multimodal contexts.
Websites: https://caiyuhu.github.io/MCiteBench
Paper: https://arxiv.org/abs/2503.02589
Code: https://github.com/caiyuhu/MCiteBench
Data Download
Please download the MCiteBench_full_dataset.zip. It contains the data.jsonl file and the visual_resources folder.
Data Statistics… See the full description on the dataset page: https://huggingface.co/datasets/caiyuhu/MCiteBench.mc4-zh-idiom-cpt
mC4 zh — Idiom-Tagged Continued-Pretraining Corpus
A 9.6M-document Chinese corpus for continued pretraining on cultural knowledge in
figurative language. Each document is natural web text (from the C4/mC4 zh subset)
containing at least one culturally meaningful chengyu, with an appended knowledge
block that lists every matched idiom together with its figurative meaning(s) and
classical source citation.
Built 2026-07-16 as Stage 1 (continue-pretraining data) of the… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/mc4-zh-idiom-cpt.terence-mckenna-transcripts
Terence McKenna Transcripts
Catalog of machine-transcribed talks and interviews, mostly by Terence McKenna, built from 172 videos. Two tables:
talks — one row per video; full original transcript (including [SPEAKER_XX] diarization tags)
turns — one row per non-empty line; speaker tags stripped from text
Speaker labels are not consistent across files (no cross-video voice fingerprinting), so turns has no speaker column.
Load
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/paulalesius/terence-mckenna-transcripts.flowjudge-dialam
FlowJudge DialAM incremental argument patches
This is the actual transformed dataset used to test whether a 0.6B open model
can learn one falsifiable behavior: given one new proposition and one complete
fixed-size block of earlier propositions from the same dialogue, emit every and
only direct SUPPORT, ATTACK, or REPHRASE edge as one bare JSON object.
An empty relation list is required when no direct edge exists.
The package includes the nested v1 data-efficiency curve, the v2… See the full description on the dataset page: https://huggingface.co/datasets/mr-mc/flowjudge-dialam.tb-explore17-mcode-m3-harness-variance
Terminal-Bench 2.1 explore-17 — mcode / MiniMax-M3 harness variance
Three complete 17-task runs of the same dataset ref with the same agent and
model, differing only in execution substrate and concurrency, plus one isolated
rerun. The point of the bundle is not the resolve rate — it is how much the
resolve rate moves when nothing about the task or the model changes.
Same everywhere: dataset ai-solution-finetune/terminal-bench-2-1-explore-17 at… See the full description on the dataset page: https://huggingface.co/datasets/miaomiao64/tb-explore17-mcode-m3-harness-variance.astro-mcq
Astro-MCQ Dataset
Astro-MCQ is the first dataset in the upcoming AstroBench collection, a suite of domain-specific benchmark datasets for evaluating small and large language models (SLMs and LLMs) in space mission engineering and astronautics.
Overview
Astro-MCQ is a multiple-choice question dataset designed to evaluate language model performance across key topics in astronautics, including:
Orbital mechanics
Space propulsion
Space environment and its effects
Spacecraft… See the full description on the dataset page: https://huggingface.co/datasets/patrickfleith/astro-mcq.NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET
Nepali Devanagari SFT Dataset — Final Clean Release
A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments.
Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns.
Dataset at a Glance
Property
Value
Total rows
100,000
Total conversation messages
200,000
Human messages
100,000
GPT messages
100,000… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET.Rationale_MCTS
Rationale MCTS Dataset: Enabling LLMs to Assess Through Rationale Thought Trees
The Rationale MCTS dataset consists of intermediate assessment rationales generated by large language models (LLMs). These rationales are "noisy," meaning they might contain errors or approximate reasoning, tailored for step-by-step explainable assessment of student answers in science and biology. The dataset targets questions from the The Hewlett Foundation: Short Answer Scoring competition, available… See the full description on the dataset page: https://huggingface.co/datasets/jiazhengli/Rationale_MCTS.stateless-mcp-agent-evaluation-suite-2026
⚡ Stateless Model Context Protocol (MCP 2026) & Agent Evaluation Suite
A Production-Grade Corpus & Evaluation Harness for Claude 5, GPT-6 Astra, and DeepSeek V4.1
⚡ Overview & Industry Problem
As of late 2026, autonomous agent engineering has superseded prompt engineering. Enterprises and developers rely on Stateless Model Context Protocol (MCP) to connect reasoning models (Claude 5 Fable, GPT-6 Astra, DeepSeek V4.1-Flash) to production backends.… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/stateless-mcp-agent-evaluation-suite-2026.mcp-server-catalog
MCP Server Catalog
A comprehensive catalog of 38 Model Context Protocol (MCP) servers for AI agents, covering data access, agent infrastructure, business-to-agent interfaces, compliance, and more.
Overview
This dataset provides a structured catalog of MCP servers that give AI agents access to real-world data and capabilities. Each server follows the MCP standard and can be used with Claude, GPT, and other LLMs that support tool use.
Categories
Category… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/mcp-server-catalog.loggenix_moe_mcs_v0
Merged Chat Dataset
Dataset Description
This dataset is a merged collection of multiple instruction-following and conversational datasets, formatted for supervised fine-tuning (SFT) of language models.
Created: 2025-08-06 08:47:50
Dataset Statistics
Total Examples: 302,417
Token Count Statistics:
Min: 50
Max: 2984
Mean: 592
Median: 473
Source Datasets
This merged dataset includes examples from the following sources:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/loggenix_moe_mcs_v0.Synthetic_Dataset_For_MCQA
