CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Vinkius /mcp-registry Vinkius Connector Registry — Open Data Initiative Welcome to the Vinkius Open Data Initiative. We are opening access to the Vinkius connector catalog. This repository provides automatically updated documentation for 9,981 unique connectors for AI agents. Research & Training Applications This highly structured corpus is designed specifically for AI researchers, data scientists, and language model developers. It provides a robust foundation for advancing artificial… See the full description on the dataset page: https://huggingface.co/datasets/Vinkius/mcp-registry.tabulartext-classification1K<n<10K4 likes859 downloads3h agoHugging Face02McAuley-Lab /Amazon-C4 Amazon-C4 A complex product search dataset built based on Amazon Reviews 2023 dataset. C4 is short for Complex Contexts Created by ChatGPT. Quick Start Loading Queries from datasets import load_dataset dataset = load_dataset('McAuley-Lab/Amazon-C4')['test'] >>> dataset Dataset({ features: ['qid', 'query', 'item_id', 'user_id', 'ori_rating', 'ori_review'], num_rows: 21223 }) >>> dataset[288] {'qid': 288, 'query': 'I need something that can entertain my… See the full description on the dataset page: https://huggingface.co/datasets/McAuley-Lab/Amazon-C4.tabular10K<n<100K8 likes331 downloads2y agoHugging Face03mcaleste /sat_multiple_choice_math_may_23This is the set of math SAT questions from the May 2023 SAT, taken from here: https://www.mcelroytutoring.com/lower.php?url=44-official-sat-pdfs-and-82-official-act-pdf-practice-tests-free. Questions that included images were not included but all other math questions, including those that have tables were included. tabularn<1K2 likes286 downloads3y agoHugging Face04fassabilf /diffusion-mcqa-gen-pelatnas-2026 Which Prompt Made This? — Generated Edition Pelatnas IOAI 2026 · Task Diffusion MCQA (varian trajectory) Sebuah model text-to-image sedang bekerja. Di tengah prosesnya, gambar belum menjadi gambar — yang ada hanya latent ter-noise: tensor 4 × 64 × 64 berisi campuran struktur yang mulai muncul dan derau Gaussian. Kali ini kalimat itu harfiah. Latent yang kamu terima benar-benar diambil dari tengah proses generate: sebuah trajectory denoising DDIM 50 langkah dihentikan sejenak… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/diffusion-mcqa-gen-pelatnas-2026.tabularimage-classificationn<1K0 likes190 downloads2mo agoHugging Face05MCINext /farsick-sts Dataset Summary FarSick STS is a Persian (Farsi) dataset designed for the Semantic Textual Similarity (STS) task. It is a part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was developed by translating and adapting the English SICK (Sentences Involving Compositional Knowledge) dataset, and it features Persian sentence pairs annotated for their degree of semantic relatedness. Language(s): Persian (Farsi) Task(s): Semantic Textual Similarity (STS) Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/farsick-sts.tabular1K<n<10K0 likes132 downloads1y agoHugging Face06Chengxiang1122 /mcl-mmcl-audiocapsaudio10K<n<100K0 likes98 downloads7mo agoHugging Face07mcemilg /turkish-plu-goal-inferenceHomepage: https://github.com/GGLAB-KU/turkish-plu tabulartext-classification100K<n<1M1 likes93 downloads3y agoHugging Face08mcemilg /turkish-plu-step-inferenceHomepage: https://github.com/GGLAB-KU/turkish-plu tabulartext-classification100K<n<1M2 likes86 downloads3y agoHugging Face09mcemilg /turkish-plu-step-orderingHomepage: https://github.com/GGLAB-KU/turkish-plu/ tabulartext-classification100K<n<1M1 likes78 downloads3y agoHugging Face10mcemilg /TrClaim19Version: v1_1 Homepage: https://github.com/YSKartal/TrClaim19 tabulartext-classification1K<n<10K0 likes76 downloads3y agoHugging Face11mcemilg /turkish-plu-next-event-predictionHomepage: https://github.com/GGLAB-KU/turkish-plu tabulartext-classification10K<n<100K1 likes74 downloads3y agoHugging Face12mcosarinsky /CheXmask-U CheXmask-U CheXmask-U is a dataset for landmark-based anatomical segmentation on chest X-ray images, providing per-node uncertainty estimates for anatomical landmarks. 📄 Paper | 💻 Code | 🌐 Project Page (Hugging Face Space) 📦 Pretrained Weights Dataset Contents The dataset is provided as CSV files. Each row corresponds to a single chest X-ray sample and contains the following fields: Image ID: Reference to the original chest X-ray image according to the source… See the full description on the dataset page: https://huggingface.co/datasets/mcosarinsky/CheXmask-U.tabularimage-segmentation100K<n<1M2 likes64 downloads9mo agoHugging Face13caiyuhu /MCiteBench MCiteBench Dataset MCiteBench is a benchmark for evaluating the ability of Multimodal Large Language Models (MLLMs) to generate text with citations in multimodal contexts. Websites: https://caiyuhu.github.io/MCiteBench Paper: https://arxiv.org/abs/2503.02589 Code: https://github.com/caiyuhu/MCiteBench Data Download Please download the MCiteBench_full_dataset.zip. It contains the data.jsonl file and the visual_resources folder. Data Statistics… See the full description on the dataset page: https://huggingface.co/datasets/caiyuhu/MCiteBench.imagetext-generation1K<n<10K0 likes62 downloads1y agoHugging Face14McGill-NLP /CHASE-QA CHASE: Challenging AI with Synthetic Evaluations The pace of evolution of Large Language Models (LLMs) necessitates new approaches for rigorous and comprehensive evaluation. Traditional human annotation is increasingly impracticable due to the complexities and costs involved in generating high-quality, challenging problems. In this work, we introduce **CHASE**, a unified framework to synthetically generate challenging problems using LLMs without human involvement. For a given task… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/CHASE-QA.imagequestion-answeringn<1K0 likes57 downloads2y agoHugging Face15crackedvibe /mcp-registry-probe-2026-09 MCP registry probe, September 2026 Row-level results of Cracked's nightly probe of every public, remote Streamable-HTTP server listed in the official Model Context Protocol registry with an open endpoint that exposed at least one tool at the last sync. This is the data behind the report The state of public MCP servers, September 2026 on cracked.ai. Method On 2026-09-02 (probe run stamped 2026-09-02T14:30:12Z), the probe connected to each server and ran, in order:… See the full description on the dataset page: https://huggingface.co/datasets/crackedvibe/mcp-registry-probe-2026-09.tabulartabular-classificationn<1K0 likes57 downloads23d agoHugging Face16AndreasXi /LP_MusicCaps_MCaudio1K<n<10K0 likes56 downloads4mo agoHugging Face17McGill-NLP /ImplicatureX ImplicatureX More information can be found at https://github.com/cesare-spinoso/ImplicatureX. import pandas as pd # skiprows=1: the first line is a leading comment, not part of the header df = pd.read_csv("implicatureX.csv", skiprows=1) Citation If you use our data, please cite us: @misc{piano2026evaluatingcommunicativebeliefupdates, title={Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation}… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/ImplicatureX.tabular1K<n<10K0 likes54 downloads2mo agoHugging Face18Databoost /MCRS_by_Databoosttabulartext-classification10K<n<100K1 likes48 downloads2y agoHugging Face19iBlessi /mcp-server-resource-benchmark MCP Server Resource Benchmark: RAM, Startup, Tool Counts Measured resident memory, startup time and tool counts for 11 popular MCP servers, plus a concurrent five-server stack. Results 178-422MB resident per server (median 188MB) 0.5-1.9 seconds warm startup 961.6MB for a concurrent five-server stack, cross-checked by two independent measurement paths (psutil and PowerShell WorkingSet64) that agreed exactly Runtime floors for context: 52MB bare Node, 15MB bare… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/mcp-server-resource-benchmark.tabularn<1K0 likes48 downloads26d agoHugging Face20lifelonglab /MCAD-CIC-3xN MCAD-CIC-3xN Dataset Summary MCAD-CIC-3xN is a multi-source continual anomaly detection benchmark scenario for network intrusion detection. It combines three CIC-family source datasets into a 13-task continual-learning scenario: CIC-IDS2017-derived tasks; CIC-IDS2018-derived tasks; CIC-UNSW-NB15-derived tasks. Unlike a single-task-per-source construction, this benchmark provides multiple concept-grouped tasks per source dataset. It is intended to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/lifelonglab/MCAD-CIC-3xN.tabulartabular-classification1M<n<10M0 likes43 downloads1mo agoHugging Face21fassabilf /diffusion-mcqa-pelatnas-2026 Which Prompt Made This? Pelatnas IOAI 2026 · Task Diffusion MCQA Sebuah model text-to-image sedang bekerja. Di tengah prosesnya, gambar belum menjadi gambar — yang ada hanya latent ter-noise: tensor 4 × 64 × 64 berisi campuran sisa struktur gambar dan derau Gaussian. Kami menangkap 250 state seperti itu. Untuk tiap state kamu tahu berapa banyak noise yang sudah ditambahkan (timestep t), dan kamu diberi 5 kandidat caption. Tepat satu adalah deskripsi asli gambarnya. Tentukan yang… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/diffusion-mcqa-pelatnas-2026.tabularimage-classificationn<1K0 likes28 downloads2mo agoHugging Face22odychlapanis /mcm-civil-procedure Dataset Card for the Multiple-Choice Mutation Civil Procedure extension data Dataset Summary The task was originally presented in the paper: @InProceedings{Bongard.et.al.2022.NLLP, title = {{The Legal Argument Reasoning Task in Civil Procedure}}, author = {Bongard, Leonard and Held, Lena and Habernal, Ivan}, booktitle = {Proceedings of the Natural Legal Language Processing Workshop 2022}, pages = {194--207}, year = {2022}… See the full description on the dataset page: https://huggingface.co/datasets/odychlapanis/mcm-civil-procedure.tabularn<1K0 likes27 downloads2y agoHugging Face23lifelonglab /MCAD-CIC-3x1 MCAD-CIC-3x1 Dataset Summary MCAD-CIC-3x1 is a multi-source continual anomaly detection benchmark scenario for network intrusion detection. It combines three CIC-family source datasets into a three-task continual-learning scenario: cicids2017 cicids2018 cicunsw Each task corresponds to one consolidated source dataset. The benchmark is designed to evaluate continual anomaly detection methods under cross-source distribution shift. The dataset contains 17,915,569… See the full description on the dataset page: https://huggingface.co/datasets/lifelonglab/MCAD-CIC-3x1.tabulartabular-classification10M<n<100M0 likes25 downloads1mo agoHugging Face24powidla /MC-III-50Welcome to the Friend or Foe Collection! tabulartabular-classification1M<n<10M0 likes20 downloads1y agoHugging Face25SINAI /MCE-Corpus Dataset Description Paper: Sentiment polarity detection in Spanish reviews combining supervised and unsupervised approaches Point of Contact: jmperea@ujaen.es, emcamara@ujaen.es MuchoCine corpus in English (MCE) is the translated version of the MuchoCine corpus (Spanish Movies Reviews). The MuchoCine corpus was developed by the researcher Fermín Cruz Mata and presented in 2008 at number 41 of the journal Natural Language Processing in the paper titled Document Classification based… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/MCE-Corpus.tabular1K<n<10K0 likes19 downloads3y agoHugging Face26qwqqwqqwq121 /Wordle-MCM-Dataset Wordle Player Performance Dataset 1. 简介 本数据集源自 2023 MCM Problem C,包含 2022 年全年的 Wordle 每日单词、报告人数及玩家猜测分布。 2. 数据内容 Date: 日期 Word: 每日目标单词 Total_Reported: 报告总人数 Hard_Mode: 硬核模式人数 Try_1 to Try_6: 玩家在第 1 至 6 次尝试中猜中的比例 Try_7_plus: 未能在 6 次内猜中的比例 3. 用途 本项目使用该数据集通过 Bi-LSTM 模型预测玩家的平均尝试步数 (Mean Tries)。 tabulartime-series-forecastingn<1K0 likes18 downloads9mo agoHugging Face27nmixx-fin /twice_kr_financial_mcqa_cls FinancialMCQA-CLS-ko Multiple-choice questions, where a question and answer choices are provided to find the correct answer. Utilizing the open dataset FINNUMBER/QA_Instruction (original source: public websites, Wikipedia). tabulartext-classification1K<n<10K0 likes15 downloads2y agoHugging Face28healthai-hq /mcp-server-grades MCP Queen: Live Grades of the MCP Ecosystem Deterministic operational grades for every remote server in the official Model Context Protocol registry, from continuous live probes by MCP Queen, the trust layer for the MCP ecosystem. 9,326 remote servers graded from 43,320+ live probes (July 2026 snapshot). Each row is a server's latest probe result. Columns column meaning server_name reverse-DNS registry name (e.g. com.healthai/clarity) title display… See the full description on the dataset page: https://huggingface.co/datasets/healthai-hq/mcp-server-grades.tabulartabular-classification1K<n<10K1 likes15 downloads2mo agoHugging Face29drelhaj /MCWC Multilingual Corpus of World’s Constitutions (MCWC) The MCWC is a curated multilingual corpus of constitutional texts from 191 countries, including both current and historical versions. The dataset provides aligned constitutional content in English, Arabic, and Spanish, enabling comparative legal analysis and multilingual NLP research. This CSV version is a cleaned, structured, and sentence-aligned representation of the corpus, suitable for machine translation, information… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/MCWC.tabular100K<n<1M0 likes13 downloads10mo agoHugging Face30mcgenxsp /CNAPStabulartext-classification100K<n<1M0 likes13 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.