CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01klieret /swe-bench-dummy-test-datasettextn<1K0 likes75k downloads1y agoHugging Face02databricks /databricks-dolly-15k Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/databricks/databricks-dolly-15k.textquestion-answering10K<n<100K1.1k likes63k downloads3y agoHugging Face03albertvillanova /datasets-tests-compressiontextn<1K0 likes60k downloads5y agoHugging Face04llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes44k downloads2y agoHugging Face05agentica-org /DeepScaleR-Preview-Dataset Data Our training dataset consists of approximately 40,000 unique mathematics problem-answer pairs compiled from: AIME (American Invitational Mathematics Examination) problems (1984-2023) AMC (American Mathematics Competition) problems (prior to 2023) Omni-MATH dataset Still dataset Format Each row in the JSON dataset contains: problem: The mathematical question text, formatted with LaTeX notation. solution: Offical solution to the problem, including LaTeX formatting… See the full description on the dataset page: https://huggingface.co/datasets/agentica-org/DeepScaleR-Preview-Dataset.text10K<n<100K206 likes39k downloads2y agoHugging Face06efficient-deep-research /synthesized_datasettext10K<n<100K0 likes34k downloads11mo agoHugging Face07common-pile /comma_v0.1_training_dataset Comma v0.1 dataset This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T. It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data. If you are looknig for the raw Common Pile v0.1 data, please see this collection. You can learn more about Common Pile in our paper. Mixing rates and token counts The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage. During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.text100M<n<1B45 likes26k downloads1y agoHugging Face08Kaphathy /Dataset MM-OphBench: Multi-Center Multimodal Clinical Ophthalmic Benchmark Dataset A Large-Scale, Standardized Multi-Center Benchmark Covering 7 Imaging Modalities & 4.3M+ Clinical Records 1. Executive Summary & Repository Overview The MM-OphBench repository hosts a petabyte-scale, clinically harmonized ophthalmic image archive compiled from leading ophthalmic hospitals and benchmark cohorts. It spans 4,307,415 high-resolution diagnostic images and multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Kaphathy/Dataset.textimage-classificationn<1K2 likes19k downloads1d agoHugging Face09zouhar /bio-mqm-datasetThis dataset is compiled from the official Amazon repository (all respective licensing applies). It contains system translations, multiple references, and their quality evaluation on the MQM scale. It accompanies the ACL 2024 paper Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains. Watch a brief 4 minutes-long video. Abstract: We introduce a new, extensive multidimensional quality metrics (MQM) annotated dataset covering 11 language pairs in the biomedical domain. We use this… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/bio-mqm-dataset.texttranslation10K<n<100K8 likes19k downloads2y agoHugging Face10RUC-NLPIR /FlashRAG_datasets ⚡FlashRAG: A Python Toolkit for Efficient RAG Research FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms. With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components. For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.textquestion-answering1M<n<10M94 likes18k downloads1y agoHugging Face11Manusagents /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M196 likes18k downloads24d agoHugging Face12KAS2003 /xfield-radar-dataset-20260915 XField radar dataset — formal snapshot, 2026-09-15 Upload status: COMPLETE — every shard verified against its remote SHA-256 and byte size. See UPLOAD_COMPLETE.json. This public research snapshot preserves the currently admitted XField/GRT simulation dataset: radar inputs, existing GT, available raw sensor products, provenance, and the frozen N141 train/validation split. It is not a claim that historical labels meet the newly repaired independent dense-GT pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/KAS2003/xfield-radar-dataset-20260915.text10K<n<100K0 likes16k downloads7d agoHugging Face13aisbergpublicorganization /telegram-news-ua-dataset Aisberg Telegram News UA A continuously updated, de-identified corpus of Ukrainian Telegram news and the discussion around it, published by the Ukrainian non-profit Aisberg (ГО «АЙЗБЕРГ»). It comes in two layers. The first is the raw monthly stream: every post from a fixed set of public news channels, with its reactions and its comment thread. The second is the analysis behind every report Aisberg publishes: posts from different channels grouped into one event, the manipulation… See the full description on the dataset page: https://huggingface.co/datasets/aisbergpublicorganization/telegram-news-ua-dataset.texttext-classification100K<n<1M4 likes16k downloads26m agoHugging Face14wenge-research /yayi2_pretrain_data 介绍/Introduction 本数据集源自雅意训练语料,我们精选了约100B数据,数据大小约为500GB。我们期望通过雅意预训练数据的开源推动中文预训练大模型开源社区的发展,并积极为此贡献力量。通过开源,我们与每一位合作伙伴共同构建雅意大模型生态。 We opensource the pre-trained dataset in this release, it should contain more than 100B tokens depending on the tokenizer you use, requiring more than 500GB of local storage. By open-sourcing the pre-trained dataset, we aim to contribute to the development of the Chinese pre-trained large language model open-source community. Through open-source, we aspire to… See the full description on the dataset page: https://huggingface.co/datasets/wenge-research/yayi2_pretrain_data.text1M<n<10M60 likes16k downloads3y agoHugging Face15JoTalbot /ua-open-data Україна: дзеркало відкритих даних (data.gov.ua) Автоматичне дзеркало публічних наборів data.gov.ua, яке підтримує пайплайн JoTalbot/ukraine. Набори Набір Файлів Джерело Єдиний державний реєстр юридичних осіб, фізичних осіб-підприємців та громадських формувань 6 — Реєстр декларацій родинних зв’язків та доброчесності 14 — Державний судновий реєстр України 9 — Публічні закупівлі на сайті Prozorro 1 — Інформація щодо стану розгляду справ 5 —… See the full description on the dataset page: https://huggingface.co/datasets/JoTalbot/ua-open-data.textn<1K1 likes12k downloads13m agoHugging Face16InternRobotics /InternData-fractal20220817_datatabular1K<n<10K1 likes11k downloads1y agoHugging Face17tascib /turkish-llm-dataset Turkish Pretraining Corpus Dataset Description This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models. This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/tascib/turkish-llm-dataset.text100M<n<1B15 likes11k downloads5mo agoHugging Face18InternRobotics /RoboInter-Data RoboInter-Data: Intermediate Representation Annotations for Robot Manipulation Rich, dense, per-frame intermediate representation annotations for robot manipulation, built on top of DROID and RH20T. Developed as part of the RoboInter project. You can try our Online demo. The annotations cover 230k episodes and include: subtasks, primitive skills, segmentation, gripper/object bounding boxes, placement proposals, affordance boxes, grasp poses, traces, contact points, etc. And each… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/RoboInter-Data.textrobotics1K<n<10K16 likes11k downloads7mo agoHugging Face19PortPy-Project /PortPy_Dataset PortPy: Planning and Optimization for Radiation Therapy Data Overview PortPy equips researchers with a robust benchmark patient dataset, sourced from the FDA-approved Eclipse commercial treatment planning system through its API. This dataset embodies all necessary elements for optimizing various machine configurations such as beam angles, aperture shapes, and leaf movements. It includes Dose Influence Matrix (AKA dose deposition matrix, dij matrix): The dose… See the full description on the dataset page: https://huggingface.co/datasets/PortPy-Project/PortPy_Dataset.textn<1K1 likes9.3k downloads5mo agoHugging Face20nvidia /Nemotron-VLM-Dataset-v2 Nemotron-VLM-Dataset v2 Versions Date Commit Changes 2025-11-05 head Fix nights_cot dataset. Fix/filter broken <think> entries. Update fintabnet instructions. Update indexes. 2025-10-28 214051e Initial Release Dataset Description Following up on Llama Nemotron VLM Dataset V1 with 3 million samples, we are releasing the Nemotron VLM Dataset V2 with almost three times as many high-quality samples. This time, our focus was on three… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-VLM-Dataset-v2.textvisual-question-answering1M<n<10M97 likes9.2k downloads9mo agoHugging Face21nvidia /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K110 likes8.3k downloads1y agoHugging Face22coref-data /knowref_60k_raw The Knowref 60K Dataset Project: https://github.com/aemami1/KnowRef60k Data source: https://github.com/aemami1/KnowRef60k/tree/28e5385d17967744ccb3bdba45fdd89d9690307d Fields annotation_strength (str): annotator agreement from 1-5 candidate_0 (str): the first candidate name candidate_1 (str): the second candidate name original_sentence (str): sentence before swapping the names swapped_sentence (str): sentence after swapping the names with square brackets marking the… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/knowref_60k_raw.text10K<n<100K0 likes7.9k downloads3y agoHugging Face23MIN-Lab /minWM-datatext1K<n<10K1 likes7.9k downloads4mo agoHugging Face24HuggingFaceH4 /instruction-datasetThis is the blind eval dataset of high-quality, diverse, human-written instructions with demonstrations. We will be using this for step 3 evaluations in our RLHF pipeline. textn<1K66 likes7.8k downloads4y agoHugging Face25johanneskirmayr /car-bench-dataset CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations. Dataset Structure The dataset is organized into task configs and mock data configs: Tasks Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of tool-call… See the full description on the dataset page: https://huggingface.co/datasets/johanneskirmayr/car-bench-dataset.tabulartext-generation1M<n<10M3 likes7.7k downloads7mo agoHugging Face26nvidia /Nemotron-Cascade-2-SFT-Data Nemotron-Cascade-2-SFT-Data We release the SFT data used for training Nemotron-Cascade-2. Data sources Math Our non-proof math prompts are sourced from Nemotron-Cascade-1-SFT and Nemotron-Math-v2, with responses generated by DeepSeek-V3.2, DeepSeek-V3.2-Speciale, and GPT-OSS-120B. For mathematical proofs, prompts are taken from Nemotron-Math-Proofs-v1 and generated using DeepSeek-V3.2-Speciale. Science We collect science prompts from… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-2-SFT-Data.text10M<n<100M75 likes7.7k downloads6mo agoHugging Face27LEMAS-Project /LEMAS-Dataset-train Overview This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project). It contains a large-scale training set (150k+ hours) and a curated evaluation set (500 utterances per language) covering 10 languages, all with word-level alignment. Fields key: unique utterance identifier; the first two characters indicate the language ID audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-train.texttext-to-speech100M<n<1B89 likes7.4k downloads6mo agoHugging Face28lfsm /ja-datasettext100K<n<1M0 likes7.2k downloads3y agoHugging Face29Weyaxi /sci-datasets Mainly science focused but other datasets exist too! Einstein models are based on this repo. text100K<n<1M28 likes7.1k downloads2y agoHugging Face30syafie-nzm /tokenized_datasettextn<1K0 likes7k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.