datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nq_open
Dataset Card for nq_open
Dataset Summary
The NQ-Open task, introduced by Lee et.al. 2019,
is an open domain question answering benchmark that is derived from Natural Questions.
The goal is to predict an English answer string for an input English question.
All questions can be answered using the contents of English Wikipedia.
Supported Tasks and Leaderboards
Open Domain Question-Answering,
EfficientQA Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/nq_open.soc-ratchakitcha
Royal Gazette Thailand (Ratchakitcha) Dataset
ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable)
โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย
Dataset Description
ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.MMFineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.DILA_OPENDATA_FR_2023
French Government Open Data (DILA) Dataset - 2023
Overview
The French Government Open Data (DILA) Dataset is a collection of text data extracted from various sources provided by the French government, specifically the Direction de l'information légale et administrative (DILA). This dataset contains a wide range of legal, administrative, and legislative documents. The data has been organized into several categories for easy access and analysis.
Dataset Splits… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/DILA_OPENDATA_FR_2023.OHR-BenchThis repository contains the OHR-Bench dataset and evaluation framework from the paper OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation.
📚 Paper | 💻 Code | 🌐 Project Page (OpenDataLab)
This repository contains the official code of OHR-Bench, a benchmark designed to evaluate the cascading impact of OCR on RAG.
News
2025.6.30: Updating the results of MongkeyOCR,
Nanonets-OCR-s and Azure Document Intelligence.
2025.6.26: OHR-Bench has been… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/OHR-Bench.MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking
MMFineReason-SFT-123K
The Hardest 7% — Less Data, More Reasoning
📖 Overview
MMFineReason-SFT-123K is a difficulty-filtered subset of MMFineReason-1.8M, containing only the hardest 7% of samples where Qwen3-VL-4B-Thinking consistently fails (pass rate = 0).
🎯 Key Highlights
123K Challenging Samples: Only instances where a 4B thinking model fails all 4 inference attemptsEfficient Training: Comparable performance to full 1.8M dataset with only 7% of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking.ODA-Math-460k
ODA-Math-460k
ODA-Math-460k is a large-scale math reasoning dataset curated from top-performing open mathematics corpora (selected via the OpenDataArena leaderboard) and refined through deduplication, benchmark decontamination, LLM-based filtering, and verifier-backed response distillation.It targets a “learnable but challenging” difficulty band: non-trivial for smaller models yet solvable by stronger reasoning models.
🧠 Dataset Summary
Domain: Mathematics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Math-460k.open_r1_dataset
Integration of publicly available datasets related to R1
We integrate all data and remove contaminated data and data with inconsistent formats. The user defaults to selecting version 'V1', with a total of 2592286 samples.
1 Relevant datasets mentioned in HuggingFace/open_r1:
(1) HuggingFaceH4/numina-deepseek-r1-qwen-7b: A dataset distilled using DeepSeek-R1-Distill-Qwen-7B.
Hugging Face downloads: 631.
(2) AI-MO/NuminaMath-TIR: A subset of 70K math-related samples… See the full description on the dataset page: https://huggingface.co/datasets/xiushenghuang/open_r1_dataset.MathLake
MathLake: A Large-Scale Mathematics Dataset
MathLake is a massive collection of 8.3 million mathematical problems aggregated from over 50 open-source datasets. Unlike datasets focused on filtering for the highest quality solutions immediately, MathLake prioritizes query comprehensiveness, serving as a universal "raw ore" for researchers to curate, distill, or annotate further. The dataset provides annotations for Difficulty, Format, and Subject, covering diverse mathematical fields… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MathLake.CiteVQA
CiteVQA
English | 简体中文
CiteVQA is a document visual question answering benchmark for faithful evidence attribution. Unlike conventional DocVQA datasets that only score the final answer, CiteVQA requires a model to answer a question with evidence grounded in the source document at the element level. The benchmark is designed to evaluate whether a system can not only answer correctly, but also cite the right supporting region in long, real-world PDFs.
The dataset contains 1,897… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/CiteVQA.open-christian-data
Open Christian Data
Open Christian Data (OCD) aims to be the single unified collection of all public domain Christian text. It exists to bring this collection together from across the internet and to structure it in a useful format for public use.
Beyond the Bible, Christian writing is poorly represented as a cohesive dataset or as data structured for AI training. This Hugging Face release is the AI-focused publication of the collection: consistent, downloadable JSON for model… See the full description on the dataset page: https://huggingface.co/datasets/OpenChristianDataOrg/open-christian-data.ODA-Fin-RL-12k
Unlocking Data Value in Finance: A Study on Distillation
and Difficulty-Aware Training
📖 Overview
ODA-Fin-RL-12K is a carefully curated dataset for reinforcement learning (RL) in financial domain, comprising 12,187 hard-but-verifiable samples. Designed to complement ODA-Fin-SFT-318K, this dataset targets challenging financial reasoning tasks with concise, reliably verifiable answers—optimized for RL training.
🎯 Key Highlights
12K Hard Samples: Curated… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Fin-RL-12k.morocco-cassation-court-decisions
Morocco Cassation Court Decisions
29,000+ full-text decisions from the Moroccan Court of Cassation (محكمة النقض)Source: juriscassation.cspj.ma — Official portal of the Supreme Council of the Judiciary (CSPJ)License: CC BY 4.0
Why this dataset exists
In 2026, accessing the jurisprudence of the Court of Cassation in Morocco requires being physically located in Morocco and armed with patience. The official website does not allow searching by date range, imposes a… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataMoroccanLaw/morocco-cassation-court-decisions.OpenDataGen-factuality-en-v0.1This synthetic dataset was generated using the Open DataGen Python library. (https://github.com/thoddnn/open-datagen)
Methodology:
Retrieve random article content from the HuggingFace Wikipedia English dataset.
Construct a Chain of Thought (CoT) to generate a Multiple Choice Question (MCQ).
Utilize a Large Language Model (LLM) to score the results then filter it.
All these steps are prompted in the 'template.json' file located in the specified code folder.
Code:… See the full description on the dataset page: https://huggingface.co/datasets/thoddnn/OpenDataGen-factuality-en-v0.1.italian-open-sft-chat-dataset
Italian Open SFT Chat Dataset
An Italian-first, model-neutral synthetic SFT and chat dataset for fine-tuning Italian-capable LLMs. It targets instruction tuning, Italian chat behavior, structured output generation, JSON/YAML/CSV format following, coding assistance, safety refusals, multi-turn dialogue and reasoning-style final answers. This v0.1.0 package does not include long-context QA records.
This dataset is intended for users searching for an Italian instruction tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/SerFabio89/italian-open-sft-chat-dataset.Open-World-Dataset
Open World Dataset (OWD)
This dataset, "Open World Dataset" (OWD), is a simple text file (.txt) containing a vast range of information and topics, covering virtually any imaginable subject. Its "open world" nature means it is not limited to a specific domain, making it extremely versatile for various natural language processing (NLP) applications.
Content
The dataset consists of plain text, with no specific formatting beyond line breaks.
Use Cases
Due to its… See the full description on the dataset page: https://huggingface.co/datasets/Immanuel-Bokkey/Open-World-Dataset.databricks-dolly-8k-qa-open-closeQR_opendata
Q&R (National Assembly and )
The database contains senators' questions with ministerial answers and questions from deputies wiht ministerial responses.
Open-ended_Questions_dialectal_data
Dataset Summary
A collection of open-ended questions that was provided to the data marathon competitors to populate KIND dataset. It was designed to elicit longer responses cultural and context-rich sentences.
For more details, please check the paper
The KIND Dataset: A Social Collaboration Approach for Nuanced Dialect Data Collection
Citation Information
@inproceedings{yamani-etal-2024-kind,
title = "The {KIND} Dataset: A Social Collaboration Approach for Nuanced… See the full description on the dataset page: https://huggingface.co/datasets/KIND-Dataset/Open-ended_Questions_dialectal_data.
