CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google-research-datasets /nq_open Dataset Card for nq_open Dataset Summary The NQ-Open task, introduced by Lee et.al. 2019, is an open domain question answering benchmark that is derived from Natural Questions. The goal is to predict an English answer string for an input English question. All questions can be answered using the contents of English Wikipedia. Supported Tasks and Leaderboards Open Domain Question-Answering, EfficientQA Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/nq_open.textquestion-answering10K<n<100K36 likes33k downloads3y agoHugging Face02open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.texttext-retrieval1M<n<10M14 likes28k downloads2h agoHugging Face03OpenDataArena /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M65 likes5.1k downloads8mo agoHugging Face04OpenDataArena /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M126 likes1.8k downloads7mo agoHugging Face05Nicolas-BZRD /DILA_OPENDATA_FR_2023 French Government Open Data (DILA) Dataset - 2023 Overview The French Government Open Data (DILA) Dataset is a collection of text data extracted from various sources provided by the French government, specifically the Direction de l'information légale et administrative (DILA). This dataset contains a wide range of legal, administrative, and legislative documents. The data has been organized into several categories for easy access and analysis. Dataset Splits… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/DILA_OPENDATA_FR_2023.texttext-classification1M<n<10M4 likes1.2k downloads3y agoHugging Face06opendatalab /OHR-BenchThis repository contains the OHR-Bench dataset and evaluation framework from the paper OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation. 📚 Paper | 💻 Code | 🌐 Project Page (OpenDataLab) This repository contains the official code of OHR-Bench, a benchmark designed to evaluate the cascading impact of OCR on RAG. News 2025.6.30: Updating the results of MongkeyOCR, Nanonets-OCR-s and Azure Document Intelligence. 2025.6.26: OHR-Bench has been… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/OHR-Bench.textquestion-answering1K<n<10K16 likes1.2k downloads1y agoHugging Face07OpenDataArena /MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking MMFineReason-SFT-123K The Hardest 7% — Less Data, More Reasoning 📖 Overview MMFineReason-SFT-123K is a difficulty-filtered subset of MMFineReason-1.8M, containing only the hardest 7% of samples where Qwen3-VL-4B-Thinking consistently fails (pass rate = 0). 🎯 Key Highlights 123K Challenging Samples: Only instances where a 4B thinking model fails all 4 inference attemptsEfficient Training: Comparable performance to full 1.8M dataset with only 7% of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking.imagevisual-question-answering100K<n<1M86 likes525 downloads8mo agoHugging Face08OpenDataArena /ODA-Math-460k ODA-Math-460k ODA-Math-460k is a large-scale math reasoning dataset curated from top-performing open mathematics corpora (selected via the OpenDataArena leaderboard) and refined through deduplication, benchmark decontamination, LLM-based filtering, and verifier-backed response distillation.It targets a “learnable but challenging” difficulty band: non-trivial for smaller models yet solvable by stronger reasoning models. 🧠 Dataset Summary Domain: Mathematics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Math-460k.tabularquestion-answering100K<n<1M105 likes463 downloads8mo agoHugging Face09xiushenghuang /open_r1_dataset Integration of publicly available datasets related to R1 We integrate all data and remove contaminated data and data with inconsistent formats. The user defaults to selecting version 'V1', with a total of 2592286 samples. 1 Relevant datasets mentioned in HuggingFace/open_r1: (1) HuggingFaceH4/numina-deepseek-r1-qwen-7b: A dataset distilled using DeepSeek-R1-Distill-Qwen-7B. Hugging Face downloads: 631. (2) AI-MO/NuminaMath-TIR: A subset of 70K math-related samples… See the full description on the dataset page: https://huggingface.co/datasets/xiushenghuang/open_r1_dataset.texttext-generation1M<n<10M5 likes409 downloads2y agoHugging Face10OpenDataArena /MathLake MathLake: A Large-Scale Mathematics Dataset MathLake is a massive collection of 8.3 million mathematical problems aggregated from over 50 open-source datasets. Unlike datasets focused on filtering for the highest quality solutions immediately, MathLake prioritizes query comprehensiveness, serving as a universal "raw ore" for researchers to curate, distill, or annotate further. The dataset provides annotations for Difficulty, Format, and Subject, covering diverse mathematical fields… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MathLake.tabularquestion-answering1M<n<10M21 likes351 downloads5mo agoHugging Face11opendatalab /CiteVQA CiteVQA English | 简体中文 CiteVQA is a document visual question answering benchmark for faithful evidence attribution. Unlike conventional DocVQA datasets that only score the final answer, CiteVQA requires a model to answer a question with evidence grounded in the source document at the element level. The benchmark is designed to evaluate whether a system can not only answer correctly, but also cite the right supporting region in long, real-world PDFs. The dataset contains 1,897… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/CiteVQA.textvisual-question-answering1K<n<10K11 likes295 downloads4mo agoHugging Face12OpenChristianDataOrg /open-christian-data Open Christian Data Open Christian Data (OCD) aims to be the single unified collection of all public domain Christian text. It exists to bring this collection together from across the internet and to structure it in a useful format for public use. Beyond the Bible, Christian writing is poorly represented as a cohesive dataset or as data structured for AI training. This Hugging Face release is the AI-focused publication of the collection: consistent, downloadable JSON for model… See the full description on the dataset page: https://huggingface.co/datasets/OpenChristianDataOrg/open-christian-data.tabulartext-generation100K<n<1M0 likes185 downloads2mo agoHugging Face13OpenDataArena /ODA-Fin-RL-12k Unlocking Data Value in Finance: A Study on Distillation and Difficulty-Aware Training 📖 Overview ODA-Fin-RL-12K is a carefully curated dataset for reinforcement learning (RL) in financial domain, comprising 12,187 hard-but-verifiable samples. Designed to complement ODA-Fin-SFT-318K, this dataset targets challenging financial reasoning tasks with concise, reliably verifiable answers—optimized for RL training. 🎯 Key Highlights 12K Hard Samples: Curated… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/ODA-Fin-RL-12k.texttext-generation10K<n<100K5 likes155 downloads7mo agoHugging Face14OpenDataMoroccanLaw /morocco-cassation-court-decisions Morocco Cassation Court Decisions 29,000+ full-text decisions from the Moroccan Court of Cassation (محكمة النقض)Source: juriscassation.cspj.ma — Official portal of the Supreme Council of the Judiciary (CSPJ)License: CC BY 4.0 Why this dataset exists In 2026, accessing the jurisprudence of the Court of Cassation in Morocco requires being physically located in Morocco and armed with patience. The official website does not allow searching by date range, imposes a… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataMoroccanLaw/morocco-cassation-court-decisions.texttext-generation10K<n<100K0 likes70 downloads2mo agoHugging Face15thoddnn /OpenDataGen-factuality-en-v0.1This synthetic dataset was generated using the Open DataGen Python library. (https://github.com/thoddnn/open-datagen) Methodology: Retrieve random article content from the HuggingFace Wikipedia English dataset. Construct a Chain of Thought (CoT) to generate a Multiple Choice Question (MCQ). Utilize a Large Language Model (LLM) to score the results then filter it. All these steps are prompted in the 'template.json' file located in the specified code folder. Code:… See the full description on the dataset page: https://huggingface.co/datasets/thoddnn/OpenDataGen-factuality-en-v0.1.textquestion-answeringn<1K1 likes64 downloads2y agoHugging Face16SerFabio89 /italian-open-sft-chat-dataset Italian Open SFT Chat Dataset An Italian-first, model-neutral synthetic SFT and chat dataset for fine-tuning Italian-capable LLMs. It targets instruction tuning, Italian chat behavior, structured output generation, JSON/YAML/CSV format following, coding assistance, safety refusals, multi-turn dialogue and reasoning-style final answers. This v0.1.0 package does not include long-context QA records. This dataset is intended for users searching for an Italian instruction tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/SerFabio89/italian-open-sft-chat-dataset.texttext-generation10K<n<100K0 likes40 downloads4mo agoHugging Face17Immanuel-Bokkey /Open-World-Dataset Open World Dataset (OWD) This dataset, "Open World Dataset" (OWD), is a simple text file (.txt) containing a vast range of information and topics, covering virtually any imaginable subject. Its "open world" nature means it is not limited to a specific domain, making it extremely versatile for various natural language processing (NLP) applications. Content The dataset consists of plain text, with no specific formatting beyond line breaks. Use Cases Due to its… See the full description on the dataset page: https://huggingface.co/datasets/Immanuel-Bokkey/Open-World-Dataset.textquestion-answering10K<n<100K0 likes18 downloads1y agoHugging Face18jtatman /databricks-dolly-8k-qa-open-closetextsummarization1K<n<10K0 likes17 downloads3y agoHugging Face19Nicolas-BZRD /QR_opendata Q&R (National Assembly and ) The database contains senators' questions with ministerial answers and questions from deputies wiht ministerial responses. textquestion-answeringn<1K0 likes15 downloads3y agoHugging Face20KIND-Dataset /Open-ended_Questions_dialectal_data Dataset Summary A collection of open-ended questions that was provided to the data marathon competitors to populate KIND dataset. It was designed to elicit longer responses cultural and context-rich sentences. For more details, please check the paper The KIND Dataset: A Social Collaboration Approach for Nuanced Dialect Data Collection Citation Information @inproceedings{yamani-etal-2024-kind, title = "The {KIND} Dataset: A Social Collaboration Approach for Nuanced… See the full description on the dataset page: https://huggingface.co/datasets/KIND-Dataset/Open-ended_Questions_dialectal_data.textquestion-answeringn<1K0 likes14 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.