datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mcp-registry
Vinkius Connector Registry — Open Data Initiative
Welcome to the Vinkius Open Data Initiative. We are opening access to the Vinkius connector catalog. This repository provides automatically updated documentation for 9,981 unique connectors for AI agents.
Research & Training Applications
This highly structured corpus is designed specifically for AI researchers, data scientists, and language model developers. It provides a robust foundation for advancing artificial… See the full description on the dataset page: https://huggingface.co/datasets/Vinkius/mcp-registry.Amazon-C4
Amazon-C4
A complex product search dataset built based on Amazon Reviews 2023 dataset.
C4 is short for Complex Contexts Created by ChatGPT.
Quick Start
Loading Queries
from datasets import load_dataset
dataset = load_dataset('McAuley-Lab/Amazon-C4')['test']
>>> dataset
Dataset({
features: ['qid', 'query', 'item_id', 'user_id', 'ori_rating', 'ori_review'],
num_rows: 21223
})
>>> dataset[288]
{'qid': 288, 'query': 'I need something that can entertain my… See the full description on the dataset page: https://huggingface.co/datasets/McAuley-Lab/Amazon-C4.sat_multiple_choice_math_may_23This is the set of math SAT questions from the May 2023 SAT, taken from here: https://www.mcelroytutoring.com/lower.php?url=44-official-sat-pdfs-and-82-official-act-pdf-practice-tests-free.
Questions that included images were not included but all other math questions, including those that have tables were included.
diffusion-mcqa-gen-pelatnas-2026
Which Prompt Made This? — Generated Edition
Pelatnas IOAI 2026 · Task Diffusion MCQA (varian trajectory)
Sebuah model text-to-image sedang bekerja. Di tengah prosesnya, gambar belum
menjadi gambar — yang ada hanya latent ter-noise: tensor 4 × 64 × 64 berisi
campuran struktur yang mulai muncul dan derau Gaussian.
Kali ini kalimat itu harfiah. Latent yang kamu terima benar-benar diambil dari
tengah proses generate: sebuah trajectory denoising DDIM 50 langkah dihentikan
sejenak… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/diffusion-mcqa-gen-pelatnas-2026.farsick-sts
Dataset Summary
FarSick STS is a Persian (Farsi) dataset designed for the Semantic Textual Similarity (STS) task. It is a part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was developed by translating and adapting the English SICK (Sentences Involving Compositional Knowledge) dataset, and it features Persian sentence pairs annotated for their degree of semantic relatedness.
Language(s): Persian (Farsi)
Task(s): Semantic Textual Similarity (STS)
Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/farsick-sts.mcl-mmcl-audiocapsturkish-plu-goal-inferenceHomepage: https://github.com/GGLAB-KU/turkish-plu
turkish-plu-step-inferenceHomepage: https://github.com/GGLAB-KU/turkish-plu
turkish-plu-step-orderingHomepage: https://github.com/GGLAB-KU/turkish-plu/
TrClaim19Version: v1_1
Homepage: https://github.com/YSKartal/TrClaim19
turkish-plu-next-event-predictionHomepage: https://github.com/GGLAB-KU/turkish-plu
CheXmask-U
CheXmask-U
CheXmask-U is a dataset for landmark-based anatomical segmentation on chest X-ray images, providing per-node uncertainty estimates for anatomical landmarks.
📄 Paper | 💻 Code | 🌐 Project Page (Hugging Face Space) 📦 Pretrained Weights
Dataset Contents
The dataset is provided as CSV files. Each row corresponds to a single chest X-ray sample and contains the following fields:
Image ID: Reference to the original chest X-ray image according to the source… See the full description on the dataset page: https://huggingface.co/datasets/mcosarinsky/CheXmask-U.MCiteBench
MCiteBench Dataset
MCiteBench is a benchmark for evaluating the ability of Multimodal Large Language Models (MLLMs) to generate text with citations in multimodal contexts.
Websites: https://caiyuhu.github.io/MCiteBench
Paper: https://arxiv.org/abs/2503.02589
Code: https://github.com/caiyuhu/MCiteBench
Data Download
Please download the MCiteBench_full_dataset.zip. It contains the data.jsonl file and the visual_resources folder.
Data Statistics… See the full description on the dataset page: https://huggingface.co/datasets/caiyuhu/MCiteBench.CHASE-QA
CHASE: Challenging AI with Synthetic Evaluations
The pace of evolution of Large Language Models (LLMs) necessitates new approaches for rigorous and comprehensive evaluation. Traditional human annotation is increasingly impracticable due to the complexities and costs involved in generating high-quality, challenging problems. In this work, we introduce **CHASE**, a unified framework to synthetically generate challenging problems using LLMs without human involvement. For a given task… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/CHASE-QA.mcp-registry-probe-2026-09
MCP registry probe, September 2026
Row-level results of Cracked's nightly probe of every public, remote Streamable-HTTP server listed in the official Model Context Protocol registry with an open endpoint that exposed at least one tool at the last sync. This is the data behind the report The state of public MCP servers, September 2026 on cracked.ai.
Method
On 2026-09-02 (probe run stamped 2026-09-02T14:30:12Z), the probe connected to each server and ran, in order:… See the full description on the dataset page: https://huggingface.co/datasets/crackedvibe/mcp-registry-probe-2026-09.LP_MusicCaps_MCImplicatureX
ImplicatureX
More information can be found at https://github.com/cesare-spinoso/ImplicatureX.
import pandas as pd
# skiprows=1: the first line is a leading comment, not part of the header
df = pd.read_csv("implicatureX.csv", skiprows=1)
Citation
If you use our data, please cite us:
@misc{piano2026evaluatingcommunicativebeliefupdates,
title={Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation}… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/ImplicatureX.MCRS_by_Databoostmcp-server-resource-benchmark
MCP Server Resource Benchmark: RAM, Startup, Tool Counts
Measured resident memory, startup time and tool counts for 11 popular MCP servers, plus a concurrent five-server stack.
Results
178-422MB resident per server (median 188MB)
0.5-1.9 seconds warm startup
961.6MB for a concurrent five-server stack, cross-checked by two independent measurement paths (psutil and PowerShell WorkingSet64) that agreed exactly
Runtime floors for context: 52MB bare Node, 15MB bare… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/mcp-server-resource-benchmark.MCAD-CIC-3xN
MCAD-CIC-3xN
Dataset Summary
MCAD-CIC-3xN is a multi-source continual anomaly detection benchmark scenario for network intrusion detection. It combines three CIC-family source datasets into a 13-task continual-learning scenario:
CIC-IDS2017-derived tasks;
CIC-IDS2018-derived tasks;
CIC-UNSW-NB15-derived tasks.
Unlike a single-task-per-source construction, this benchmark provides multiple concept-grouped tasks per source dataset. It is intended to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/lifelonglab/MCAD-CIC-3xN.diffusion-mcqa-pelatnas-2026
Which Prompt Made This?
Pelatnas IOAI 2026 · Task Diffusion MCQA
Sebuah model text-to-image sedang bekerja. Di tengah prosesnya, gambar belum
menjadi gambar — yang ada hanya latent ter-noise: tensor 4 × 64 × 64 berisi
campuran sisa struktur gambar dan derau Gaussian.
Kami menangkap 250 state seperti itu. Untuk tiap state kamu tahu berapa banyak
noise yang sudah ditambahkan (timestep t), dan kamu diberi 5 kandidat
caption. Tepat satu adalah deskripsi asli gambarnya.
Tentukan yang… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/diffusion-mcqa-pelatnas-2026.mcm-civil-procedure
Dataset Card for the Multiple-Choice Mutation Civil Procedure extension data
Dataset Summary
The task was originally presented in the paper:
@InProceedings{Bongard.et.al.2022.NLLP,
title = {{The Legal Argument Reasoning Task in Civil Procedure}},
author = {Bongard, Leonard and Held, Lena and Habernal, Ivan},
booktitle = {Proceedings of the Natural Legal Language Processing
Workshop 2022},
pages = {194--207},
year = {2022}… See the full description on the dataset page: https://huggingface.co/datasets/odychlapanis/mcm-civil-procedure.MCAD-CIC-3x1
MCAD-CIC-3x1
Dataset Summary
MCAD-CIC-3x1 is a multi-source continual anomaly detection benchmark scenario for network intrusion detection. It combines three CIC-family source datasets into a three-task continual-learning scenario:
cicids2017
cicids2018
cicunsw
Each task corresponds to one consolidated source dataset. The benchmark is designed to evaluate continual anomaly detection methods under cross-source distribution shift.
The dataset contains 17,915,569… See the full description on the dataset page: https://huggingface.co/datasets/lifelonglab/MCAD-CIC-3x1.MC-III-50Welcome to the Friend or Foe Collection!
MCE-Corpus
Dataset Description
Paper: Sentiment polarity detection in Spanish reviews combining supervised and unsupervised approaches
Point of Contact: jmperea@ujaen.es, emcamara@ujaen.es
MuchoCine corpus in English (MCE) is the translated version of the MuchoCine corpus (Spanish Movies Reviews). The MuchoCine corpus was developed by the researcher Fermín Cruz Mata and presented in 2008 at number 41 of the journal Natural Language Processing in the paper titled Document Classification based… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/MCE-Corpus.Wordle-MCM-Dataset
Wordle Player Performance Dataset
1. 简介
本数据集源自 2023 MCM Problem C,包含 2022 年全年的 Wordle 每日单词、报告人数及玩家猜测分布。
2. 数据内容
Date: 日期
Word: 每日目标单词
Total_Reported: 报告总人数
Hard_Mode: 硬核模式人数
Try_1 to Try_6: 玩家在第 1 至 6 次尝试中猜中的比例
Try_7_plus: 未能在 6 次内猜中的比例
3. 用途
本项目使用该数据集通过 Bi-LSTM 模型预测玩家的平均尝试步数 (Mean Tries)。
twice_kr_financial_mcqa_cls
FinancialMCQA-CLS-ko
Multiple-choice questions, where a question and answer choices are provided to find the correct answer.
Utilizing the open dataset FINNUMBER/QA_Instruction (original source: public websites, Wikipedia).
mcp-server-grades
MCP Queen: Live Grades of the MCP Ecosystem
Deterministic operational grades for every remote server in the official
Model Context Protocol registry, from continuous live probes by
MCP Queen, the trust layer for the MCP ecosystem.
9,326 remote servers graded from 43,320+ live probes (July 2026
snapshot). Each row is a server's latest probe result.
Columns
column
meaning
server_name
reverse-DNS registry name (e.g. com.healthai/clarity)
title
display… See the full description on the dataset page: https://huggingface.co/datasets/healthai-hq/mcp-server-grades.MCWC
Multilingual Corpus of World’s Constitutions (MCWC)
The MCWC is a curated multilingual corpus of constitutional texts from 191 countries, including both current and historical versions. The dataset provides aligned constitutional content in English, Arabic, and Spanish, enabling comparative legal analysis and multilingual NLP research.
This CSV version is a cleaned, structured, and sentence-aligned representation of the corpus, suitable for machine translation, information… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/MCWC.CNAPS
