datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Shamela4_Full_DB
Shamela 4 — Full Islamic Library Corpus
A complete extraction of al-Maktaba al-Shamela (الشاملة) v4, containing 8,589 books across 40 categories of classical Islamic sciences. Extracted from the original Lucene + Sqlite Shamela DB on 2026-04-26 with ~7.6 million pages and ~19 GB of Arabic text.
Dataset Structure
stage0_raw/
├── _meta/ # Cross-cutting metadata (Parquet + JSONL)
│ ├── extraction_manifest.json # Global extraction record
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AuthenticIlm/Shamela4_Full_DB.PromptEval_MMLU_full
MMLU Multi-Prompt Evaluation Data
Overview
This dataset contains the results of a comprehensive evaluation of various Large Language Models (LLMs) using multiple prompt templates on the Massive Multitask Language Understanding (MMLU) benchmark. The data is introduced in
Maia Polo, Felipe, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. "Efficient multi-prompt evaluation of LLMs."… See the full description on the dataset page: https://huggingface.co/datasets/PromptEval/PromptEval_MMLU_full.MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ericktwo/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.WildChat-4.8M-Full
Dataset Card for WildChat-4.8M-Full
Dataset Description
Interactive Search Tool: https://wildvisualizer.com
WildChat paper: https://arxiv.org/abs/2405.01470
WildVis paper: https://arxiv.org/abs/2409.03753
Point of Contact: Yuntian Deng
Dataset Summary
WildChat-4.8M-Full is a collection of 4,743,336 conversations (out of 4,804,190 originally, after removing all conversations flagged with "sexual/minors" by OpenAI Moderation) between human users and… See the full description on the dataset page: https://huggingface.co/datasets/yuntian-deng/WildChat-4.8M-Full.ot-full
Open Telco Full Benchmarks
20,588 telecom-specific evaluation samples across 8 benchmarks — the complete evaluation suite for measuring telecom AI performance.
Use this dataset for final, publishable results. For fast iteration during model development, use GSMA/ot-lite.
Eval Framework | Sample Data
Benchmarks
| Config | Samples | Task | Paper |
|--------|--------:|------|-------|
| teleqna | 10,000 | Multiple-choice Q&A on telecom standards | arXiv |
| teletables |… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/ot-full.DeepSWEGym2-Full
Dataset Description
This dataset is a merge containing many high quality SWE datasets, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills.
It is NOT specifically filtered for rows with complex/long code problems in the original datasets, despite still having an average row size of 169.08kb, a total uncompressed size of 19.80GB, and a total of 122791 examples.
Dataset Details
Curated by: MoreThought
Funded by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym2-Full.DeepSWEGym-Full
Dataset Description
This dataset is a merged version of all the SWE-bench/SWE-smith-lang datasets (88k rows total), it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills.
It is NOT specifically filtered for rows with complex/long code problems in the original datasets, despite still having an average row size of 85.4kb, a total uncompressed size of 7.53GB, and 88130 examples total.
Dataset Details
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym-Full.conseil-detat-full-documents
Décisions du Conseil d'État (France)
Description
Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé.
Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité.
Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.conseil-administratives-appel-full-documents
Décisions de Justice Administrative Françaises
Description du dataset
Ce dataset contient un corpus de décisions de justice administrative françaises extraites de la plateforme officielle Open Data de la Justice Administrative. Les documents sont publiés par le Conseil d'État dans le cadre de la politique d'ouverture des données publiques de la justice française. Les décisions sont diffusées au format XML et anonymisées conformément aux exigences légales relatives… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-administratives-appel-full-documents.Arabic-VLM-Full-Pearl
💎 The Arabic VLM Dataset (Full Pearl Edition)
This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper.
Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.service_public_part-full-documents
🇫🇷 Dataset Service-Public.fr – Fiches administratives structurées
Ce dataset est constitué à partir des contenus officiels publiés sur la plateformeService-Public.fr.Il regroupe des fiches pratiques et ressources administratives à destination des particuliers et des professionnels, couvrant un large éventail de démarches et de thématiques de l’administration française.
La structure et la méthodologie de ce dataset sont fortement inspirées du dataset Service-Public.fr practical… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/service_public_part-full-documents.service_public_pro-full-documentstravail_emploi-full-documents
🇫🇷 Dataset Ministère du Travail et de l’Emploi – Fiches structurées
Ce dataset est constitué à partir des contenus publics diffusés sur le site officiel duMinistère du Travail et de l’Emploi :https://travail-emploi.gouv.fr/
Les données sources proviennent du dépôt GitHub officiel de l’administration française :https://github.com/SocialGouv/fiches-travail-data
La structure et la logique générale de ce dataset sont inspirées du dataset Travail Emploi website Dataset, publié sur… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/travail_emploi-full-documents.gaze_dataset_full
gaze_dataset_full — StreamGaze_v2 + EgoGazeVQA + HD-EPIC
A single repository containing three complementary benchmarks for
evaluating multimodal LLMs on gaze-grounded egocentric video question
answering:
Subfolder
Source
Held-out (val/test)
Questions
StreamGaze_v2/
egoexolearn, holoassist, egtea
egtea
8 MCQ tasks (4-opt) — gaze-conditioned past/present/future
EgoGazeVQA/
ego4d, egoexo, egtea
egtea
causal / spatial / temporal (5-opt)
HD-EPIC/
P01–P09
P09… See the full description on the dataset page: https://huggingface.co/datasets/Peanuttoad/gaze_dataset_full.flan2021-full
Task Name
FLAN-2021 -> 70
{
"ag_news_subset": 108497,
"ai2_arc/ARC-Challenge": 829,
"ai2_arc/ARC-Easy": 1927,
"aeslc": 13187,
"anli/r1": 15361,
"anli/r2": 41133,
"anli/r3": 91048,
"bool_q": 8343,
"cnn_dailymail": 259607,
"coqa": 6456,
"cosmos_qa": 22996,
"definite_pronoun_resolution": 1079,
"drop": 70045,
"fix_punct": 25690,
"gem/common_gen": 60936,
"gem/dart": 56724,
"gem/e2e_nlg": 30337,
"gem/web_nlg_en": 31899… See the full description on the dataset page: https://huggingface.co/datasets/aslawliet/flan2021-full.fully-open-meditron
Fully Open Meditron Corpus
👋 Join our LiGHT community.
📖 Check out the MeditronFO blog and MeditronFO preprint.
🔜 If you are a clinician join the MOOVE initiative here.
[Hugging Face]
[Preprint]
[GitHub]
[Dataset]
License: Apache 2.0 | Authors: LiGHT
[!Note]
A clinician-vetted training corpus for medical large language models, accompanying the paper Fully Open Meditron: An Auditable Pipeline for Clinical LLMs.
The… See the full description on the dataset page: https://huggingface.co/datasets/EPFLiGHT/fully-open-meditron.tribunal-administratif-full-documents
Tribunal Administratif - Full Documents
Description
Ce dataset contient un corpus de décisions issues des Tribunaux Administratifs français, converties en documents textuels exploitables pour les applications d'intelligence artificielle.
L'objectif est de fournir un corpus prêt à l'emploi pour :
le Retrieval-Augmented Generation (RAG) ;
la recherche juridique ;
la question-réponse ;
la classification documentaire ;
le fine-tuning de modèles de langage spécialisés… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/tribunal-administratif-full-documents.math_full_minus_math500
MATH (minus MATH-500)
This dataset is derived from the original MATH dataset by Hendrycks et al.
(qwedsacf/competition_math) with all problems from the MATH-500 benchmark set removed.
Construction
Source: 12,500 problems from the MATH dataset by Hendrycks et al. (qwedsacf/competition_math)
Benchmark held out: 500 problems from the MATH-500 dataset (HuggingFaceH4/MATH-500)
Matching criterion: exact match on the problem field (see… See the full description on the dataset page: https://huggingface.co/datasets/rasbt/math_full_minus_math500.rag-qa-fulltext-ptbr
RAG QA Full-Text PT-BR Mistral
A large-scale dataset of Brazilian Portuguese RAG-style question-answer pairs
with grounded evidence spans, generated from Madras1/corpus-ptbr-v1 documents
using Mistral models. Every answer is anchored to literal quotations from the
source text, making this dataset suitable for training and evaluating
retrieval-augmented generation systems, extractive QA models, and reading
comprehension benchmarks in Portuguese.
Two configurations are available:… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/rag-qa-fulltext-ptbr.NuminaMath-full
Merged Math Datasets (Full)
This dataset combines multiple mathematical datasets for training and evaluation purposes. This version contains the complete training dataset from all source datasets.
Dataset Description
A comprehensive collection of mathematical problems and solutions from various sources, organized into training and multiple test subsets.
Dataset Structure
Training Set
Size: 217389 examples
Fields: source, question, answer
Sources:… See the full description on the dataset page: https://huggingface.co/datasets/weijiezz/NuminaMath-full.unpredictable_fullThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.rag-full-20000
Retrieval-Augmented Generation (RAG) Full 20000
Retrieval-Augmented Generation (RAG) Full 20000 is an English dataset designed for RAG-optimized models, built by Neural Bridge AI, and released under Apache license 2.0.
Dataset Description
Dataset Summary
Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by allowing them to consult an external authoritative knowledge base before generating responses. This approach significantly boosts… See the full description on the dataset page: https://huggingface.co/datasets/neural-bridge/rag-full-20000.creativemath_fullFull-Ecom-Chatbot-Dataset
E-commerce Chatbot Training Data
A curated, multi-source dataset for training and evaluating e-commerce conversational AI systems. It covers a broad range of customer intents — from product discovery and order management to returns, tool-augmented responses, and RAG-grounded Q&A — across 16+ product domains.
Dataset Summary
Split
Records
Train
35,213
Test
8,818
Total
44,031
The train/test split uses prompt-group-level stratified sampling on source ×… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Full-Ecom-Chatbot-Dataset.browsecomp-plus-structured-fullThis dataset is associated with the paper Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents.
Repository: ielab/skim-search-agent
Citation
@misc{wang2026sieve,
title = {Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents},
author = {Wang, Shuai and Chen, Haodong and Yin, Yu and Zhuang, Shengyao and
Koopman, Bevan and Zuccon, Guido},
year = {2026},
eprint =… See the full description on the dataset page: https://huggingface.co/datasets/wshuai190/browsecomp-plus-structured-full.WildChat-1M-Full
Dataset Card for WildChat-1M-Full
Dataset Description
Paper: https://arxiv.org/abs/2405.01470
Interactive Search Tool: https://wildvisualizer.com (paper)
License: ODC-BY
Language(s) (NLP): multi-lingual
Point of Contact: Yuntian Deng
Dataset Summary
WildChat-1M-Full is a collection of 1 million conversations between human users and ChatGPT, alongside demographic data, including state, country, hashed IP addresses, and request headers. We collected… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildChat-1M-Full.WildChat-4.8M-Full
Dataset Card for WildChat-4.8M-Full
Dataset Description
Interactive Search Tool: https://wildvisualizer.com
WildChat paper: https://arxiv.org/abs/2405.01470
WildVis paper: https://arxiv.org/abs/2409.03753
Point of Contact: Yuntian Deng
Dataset Summary
WildChat-4.8M-Full is a collection of 4,743,336 conversations (out of 4,804,190 originally, after removing all conversations flagged with "sexual/minors" by OpenAI Moderation) between human users and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildChat-4.8M-Full.full-modality-data
Full Modality Dataset Statistics
Video Statistics
Total Videos: 28,472
Total Duration: 1422.33 hours
Average Duration: 179.84 seconds
Median Duration: 160.08 seconds
Duration Range: 10.04s - 1780.03s
QA Statistics
Total Questions: 1,444,526
Average Questions per Video: 50.7
Questions per Video Range: 14 - 450
Question Type Distribution
OE: 1,444,526 (100.0%)
Question Category Distribution
temporal: 96,873 (6.7%)
causal: 96,873… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/full-modality-data.CHSA-Triage-Medic-Full-Dataset
CHSA-Triage-Medic-Full-Dataset
Ce dataset a été constitué dans le cadre d'un projet de formation AI Engineer (Projet CHSA). Il est conçu pour entraîner un Assistant Médical Intelligent capable d'effectuer du triage d'urgence et de fournir des raisonnements cliniques.
Le dataset est divisé en 3 sous-ensembles distincts correspondant aux différentes phases d'entraînement (Fine-Tuning Supervisé et Alignement).
Organisation du Dataset
Le repository contient trois… See the full description on the dataset page: https://huggingface.co/datasets/cyrille-elie/CHSA-Triage-Medic-Full-Dataset.
