CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AuthenticIlm /Shamela4_Full_DB Shamela 4 — Full Islamic Library Corpus A complete extraction of al-Maktaba al-Shamela (الشاملة) v4, containing 8,589 books across 40 categories of classical Islamic sciences. Extracted from the original Lucene + Sqlite Shamela DB on 2026-04-26 with ~7.6 million pages and ~19 GB of Arabic text. Dataset Structure stage0_raw/ ├── _meta/ # Cross-cutting metadata (Parquet + JSONL) │ ├── extraction_manifest.json # Global extraction record │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AuthenticIlm/Shamela4_Full_DB.text-generation10M<n<100M29 likes11k downloads4mo agoHugging Face02PromptEval /PromptEval_MMLU_full MMLU Multi-Prompt Evaluation Data Overview This dataset contains the results of a comprehensive evaluation of various Large Language Models (LLMs) using multiple prompt templates on the Massive Multitask Language Understanding (MMLU) benchmark. The data is introduced in Maia Polo, Felipe, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. "Efficient multi-prompt evaluation of LLMs."… See the full description on the dataset page: https://huggingface.co/datasets/PromptEval/PromptEval_MMLU_full.tabularquestion-answering10M<n<100M3 likes8.2k downloads2y agoHugging Face03OpenDataArena /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M65 likes5.3k downloads8mo agoHugging Face04ericktwo /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ericktwo/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M1 likes1.4k downloads8mo agoHugging Face05yuntian-deng /WildChat-4.8M-Fullgated Dataset Card for WildChat-4.8M-Full Dataset Description Interactive Search Tool: https://wildvisualizer.com WildChat paper: https://arxiv.org/abs/2405.01470 WildVis paper: https://arxiv.org/abs/2409.03753 Point of Contact: Yuntian Deng Dataset Summary WildChat-4.8M-Full is a collection of 4,743,336 conversations (out of 4,804,190 originally, after removing all conversations flagged with "sexual/minors" by OpenAI Moderation) between human users and… See the full description on the dataset page: https://huggingface.co/datasets/yuntian-deng/WildChat-4.8M-Full.texttext-generation1M<n<10M10 likes881 downloads1y agoHugging Face06GSMA /ot-full Open Telco Full Benchmarks 20,588 telecom-specific evaluation samples across 8 benchmarks — the complete evaluation suite for measuring telecom AI performance. Use this dataset for final, publishable results. For fast iteration during model development, use GSMA/ot-lite. Eval Framework | Sample Data Benchmarks | Config | Samples | Task | Paper | |--------|--------:|------|-------| | teleqna | 10,000 | Multiple-choice Q&A on telecom standards | arXiv | | teletables |… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/ot-full.textquestion-answering10K<n<100K7 likes868 downloads6mo agoHugging Face07MoreThought /DeepSWEGym2-Full Dataset Description This dataset is a merge containing many high quality SWE datasets, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is NOT specifically filtered for rows with complex/long code problems in the original datasets, despite still having an average row size of 169.08kb, a total uncompressed size of 19.80GB, and a total of 122791 examples. Dataset Details Curated by: MoreThought Funded by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym2-Full.texttext-generation100K<n<1M1 likes829 downloads16d agoHugging Face08MoreThought /DeepSWEGym-Full Dataset Description This dataset is a merged version of all the SWE-bench/SWE-smith-lang datasets (88k rows total), it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is NOT specifically filtered for rows with complex/long code problems in the original datasets, despite still having an average row size of 85.4kb, a total uncompressed size of 7.53GB, and 88130 examples total. Dataset Details Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym-Full.texttext-generation10K<n<100K1 likes680 downloads18d agoHugging Face09hulk10 /conseil-detat-full-documents Décisions du Conseil d'État (France) Description Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé. Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité. Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.tabularquestion-answering1K<n<10K1 likes490 downloads18h agoHugging Face10hulk10 /conseil-administratives-appel-full-documents Décisions de Justice Administrative Françaises Description du dataset Ce dataset contient un corpus de décisions de justice administrative françaises extraites de la plateforme officielle Open Data de la Justice Administrative. Les documents sont publiés par le Conseil d'État dans le cadre de la politique d'ouverture des données publiques de la justice française. Les décisions sont diffusées au format XML et anonymisées conformément aux exigences légales relatives… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-administratives-appel-full-documents.texttext-classification1K<n<10K1 likes486 downloads18h agoHugging Face11MohamedRashad /Arabic-VLM-Full-Pearl 💎 The Arabic VLM Dataset (Full Pearl Edition) This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper. Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.imagequestion-answering100K<n<1M10 likes449 downloads10mo agoHugging Face12hulk10 /service_public_part-full-documents 🇫🇷 Dataset Service-Public.fr – Fiches administratives structurées Ce dataset est constitué à partir des contenus officiels publiés sur la plateformeService-Public.fr.Il regroupe des fiches pratiques et ressources administratives à destination des particuliers et des professionnels, couvrant un large éventail de démarches et de thématiques de l’administration française. La structure et la méthodologie de ce dataset sont fortement inspirées du dataset Service-Public.fr practical… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/service_public_part-full-documents.textquestion-answering1K<n<10K0 likes389 downloads18h agoHugging Face13hulk10 /service_public_pro-full-documentstextquestion-answering1K<n<10K0 likes375 downloads18h agoHugging Face14hulk10 /travail_emploi-full-documents 🇫🇷 Dataset Ministère du Travail et de l’Emploi – Fiches structurées Ce dataset est constitué à partir des contenus publics diffusés sur le site officiel duMinistère du Travail et de l’Emploi :https://travail-emploi.gouv.fr/ Les données sources proviennent du dépôt GitHub officiel de l’administration française :https://github.com/SocialGouv/fiches-travail-data La structure et la logique générale de ce dataset sont inspirées du dataset Travail Emploi website Dataset, publié sur… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/travail_emploi-full-documents.textquestion-answeringn<1K0 likes357 downloads18h agoHugging Face15Peanuttoad /gaze_dataset_full gaze_dataset_full — StreamGaze_v2 + EgoGazeVQA + HD-EPIC A single repository containing three complementary benchmarks for evaluating multimodal LLMs on gaze-grounded egocentric video question answering: Subfolder Source Held-out (val/test) Questions StreamGaze_v2/ egoexolearn, holoassist, egtea egtea 8 MCQ tasks (4-opt) — gaze-conditioned past/present/future EgoGazeVQA/ ego4d, egoexo, egtea egtea causal / spatial / temporal (5-opt) HD-EPIC/ P01–P09 P09… See the full description on the dataset page: https://huggingface.co/datasets/Peanuttoad/gaze_dataset_full.question-answering0 likes348 downloads2mo agoHugging Face16aslawliet /flan2021-full Task Name FLAN-2021 -> 70 { "ag_news_subset": 108497, "ai2_arc/ARC-Challenge": 829, "ai2_arc/ARC-Easy": 1927, "aeslc": 13187, "anli/r1": 15361, "anli/r2": 41133, "anli/r3": 91048, "bool_q": 8343, "cnn_dailymail": 259607, "coqa": 6456, "cosmos_qa": 22996, "definite_pronoun_resolution": 1079, "drop": 70045, "fix_punct": 25690, "gem/common_gen": 60936, "gem/dart": 56724, "gem/e2e_nlg": 30337, "gem/web_nlg_en": 31899… See the full description on the dataset page: https://huggingface.co/datasets/aslawliet/flan2021-full.texttext-generation10M<n<100M2 likes330 downloads2y agoHugging Face17EPFLiGHT /fully-open-meditron Fully Open Meditron Corpus 👋 Join our LiGHT community. 📖 Check out the MeditronFO blog and MeditronFO preprint. 🔜 If you are a clinician join the MOOVE initiative here. [Hugging Face] [Preprint] [GitHub] [Dataset] License: Apache 2.0 | Authors: LiGHT [!Note] A clinician-vetted training corpus for medical large language models, accompanying the paper Fully Open Meditron: An Auditable Pipeline for Clinical LLMs. The… See the full description on the dataset page: https://huggingface.co/datasets/EPFLiGHT/fully-open-meditron.textquestion-answering100K<n<1M8 likes329 downloads3mo agoHugging Face18hulk10 /tribunal-administratif-full-documents Tribunal Administratif - Full Documents Description Ce dataset contient un corpus de décisions issues des Tribunaux Administratifs français, converties en documents textuels exploitables pour les applications d'intelligence artificielle. L'objectif est de fournir un corpus prêt à l'emploi pour : le Retrieval-Augmented Generation (RAG) ; la recherche juridique ; la question-réponse ; la classification documentaire ; le fine-tuning de modèles de langage spécialisés… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/tribunal-administratif-full-documents.textquestion-answering10K<n<100K0 likes288 downloads2d agoHugging Face19rasbt /math_full_minus_math500 MATH (minus MATH-500) This dataset is derived from the original MATH dataset by Hendrycks et al. (qwedsacf/competition_math) with all problems from the MATH-500 benchmark set removed.   Construction Source: 12,500 problems from the MATH dataset by Hendrycks et al. (qwedsacf/competition_math) Benchmark held out: 500 problems from the MATH-500 dataset (HuggingFaceH4/MATH-500) Matching criterion: exact match on the problem field (see… See the full description on the dataset page: https://huggingface.co/datasets/rasbt/math_full_minus_math500.texttext-generation10K<n<100K2 likes168 downloads9mo agoHugging Face20Madras1 /rag-qa-fulltext-ptbr RAG QA Full-Text PT-BR Mistral A large-scale dataset of Brazilian Portuguese RAG-style question-answer pairs with grounded evidence spans, generated from Madras1/corpus-ptbr-v1 documents using Mistral models. Every answer is anchored to literal quotations from the source text, making this dataset suitable for training and evaluating retrieval-augmented generation systems, extractive QA models, and reading comprehension benchmarks in Portuguese. Two configurations are available:… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/rag-qa-fulltext-ptbr.tabularquestion-answering1M<n<10M0 likes168 downloads5mo agoHugging Face21weijiezz /NuminaMath-full Merged Math Datasets (Full) This dataset combines multiple mathematical datasets for training and evaluation purposes. This version contains the complete training dataset from all source datasets. Dataset Description A comprehensive collection of mathematical problems and solutions from various sources, organized into training and multiple test subsets. Dataset Structure Training Set Size: 217389 examples Fields: source, question, answer Sources:… See the full description on the dataset page: https://huggingface.co/datasets/weijiezz/NuminaMath-full.textquestion-answering100K<n<1M0 likes148 downloads1y agoHugging Face22MicPie /unpredictable_fullThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice1M<n<10M3 likes147 downloads4y agoHugging Face23neural-bridge /rag-full-20000 Retrieval-Augmented Generation (RAG) Full 20000 Retrieval-Augmented Generation (RAG) Full 20000 is an English dataset designed for RAG-optimized models, built by Neural Bridge AI, and released under Apache license 2.0. Dataset Description Dataset Summary Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by allowing them to consult an external authoritative knowledge base before generating responses. This approach significantly boosts… See the full description on the dataset page: https://huggingface.co/datasets/neural-bridge/rag-full-20000.textquestion-answering10K<n<100K22 likes128 downloads3y agoHugging Face24windchimeran /creativemath_fulltabularquestion-answering1K<n<10K0 likes124 downloads1y agoHugging Face25rescommons /Full-Ecom-Chatbot-Dataset E-commerce Chatbot Training Data A curated, multi-source dataset for training and evaluating e-commerce conversational AI systems. It covers a broad range of customer intents — from product discovery and order management to returns, tool-augmented responses, and RAG-grounded Q&A — across 16+ product domains. Dataset Summary Split Records Train 35,213 Test 8,818 Total 44,031 The train/test split uses prompt-group-level stratified sampling on source ×… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Full-Ecom-Chatbot-Dataset.tabularquestion-answering10K<n<100K0 likes114 downloads6mo agoHugging Face26wshuai190 /browsecomp-plus-structured-fullThis dataset is associated with the paper Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents. Repository: ielab/skim-search-agent Citation @misc{wang2026sieve, title = {Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents}, author = {Wang, Shuai and Chen, Haodong and Yin, Yu and Zhuang, Shengyao and Koopman, Bevan and Zuccon, Guido}, year = {2026}, eprint =… See the full description on the dataset page: https://huggingface.co/datasets/wshuai190/browsecomp-plus-structured-full.question-answering0 likes106 downloads2mo agoHugging Face27allenai /WildChat-1M-Fullgated Dataset Card for WildChat-1M-Full Dataset Description Paper: https://arxiv.org/abs/2405.01470 Interactive Search Tool: https://wildvisualizer.com (paper) License: ODC-BY Language(s) (NLP): multi-lingual Point of Contact: Yuntian Deng Dataset Summary WildChat-1M-Full is a collection of 1 million conversations between human users and ChatGPT, alongside demographic data, including state, country, hashed IP addresses, and request headers. We collected… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildChat-1M-Full.texttext-generation100K<n<1M40 likes105 downloads2y agoHugging Face28allenai /WildChat-4.8M-Fullgated Dataset Card for WildChat-4.8M-Full Dataset Description Interactive Search Tool: https://wildvisualizer.com WildChat paper: https://arxiv.org/abs/2405.01470 WildVis paper: https://arxiv.org/abs/2409.03753 Point of Contact: Yuntian Deng Dataset Summary WildChat-4.8M-Full is a collection of 4,743,336 conversations (out of 4,804,190 originally, after removing all conversations flagged with "sexual/minors" by OpenAI Moderation) between human users and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildChat-4.8M-Full.texttext-generation1M<n<10M38 likes75 downloads1y agoHugging Face29lmms-lab /full-modality-data Full Modality Dataset Statistics Video Statistics Total Videos: 28,472 Total Duration: 1422.33 hours Average Duration: 179.84 seconds Median Duration: 160.08 seconds Duration Range: 10.04s - 1780.03s QA Statistics Total Questions: 1,444,526 Average Questions per Video: 50.7 Questions per Video Range: 14 - 450 Question Type Distribution OE: 1,444,526 (100.0%) Question Category Distribution temporal: 96,873 (6.7%) causal: 96,873… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/full-modality-data.tabularquestion-answering1M<n<10M1 likes72 downloads1y agoHugging Face30cyrille-elie /CHSA-Triage-Medic-Full-Dataset CHSA-Triage-Medic-Full-Dataset Ce dataset a été constitué dans le cadre d'un projet de formation AI Engineer (Projet CHSA). Il est conçu pour entraîner un Assistant Médical Intelligent capable d'effectuer du triage d'urgence et de fournir des raisonnements cliniques. Le dataset est divisé en 3 sous-ensembles distincts correspondant aux différentes phases d'entraînement (Fine-Tuning Supervisé et Alignement). Organisation du Dataset Le repository contient trois… See the full description on the dataset page: https://huggingface.co/datasets/cyrille-elie/CHSA-Triage-Medic-Full-Dataset.texttext-generation10K<n<100K0 likes61 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.