datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Darwin-CCswe-bench-dummy-test-datasetdatabricks-dolly-15k
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification,
closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/databricks/databricks-dolly-15k.datasets-tests-compressiontiny-supervised-datasetDeepScaleR-Preview-Dataset
Data
Our training dataset consists of approximately 40,000 unique mathematics problem-answer pairs compiled from:
AIME (American Invitational Mathematics Examination) problems (1984-2023)
AMC (American Mathematics Competition) problems (prior to 2023)
Omni-MATH dataset
Still dataset
Format
Each row in the JSON dataset contains:
problem: The mathematical question text, formatted with LaTeX notation.
solution: Offical solution to the problem, including LaTeX formatting… See the full description on the dataset page: https://huggingface.co/datasets/agentica-org/DeepScaleR-Preview-Dataset.synthesized_datasetDABstep
Data Agent Benchmark for Multi-step Reasoning (DABstep) Dataset
This repository hosts a HF Dataset the supports the benchmark and leaderboard.
For the main entrypoint to the benchmark, see the leaderboard here:
https://huggingface.co/spaces/adyen/DABstep
This Dataset has 3 splits:
tasks
submissions
task_scores
Users of the benchmark would read from the tasks split to run the baseline. The other splits are used to support the leaderboard.
The datasets are in the data/context… See the full description on the dataset page: https://huggingface.co/datasets/adyen/DABstep.soc-ratchakitcha
Royal Gazette Thailand (Ratchakitcha) Dataset
ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable)
โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย
Dataset Description
ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.comma_v0.1_training_dataset
Comma v0.1 dataset
This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T.
It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data.
If you are looknig for the raw Common Pile v0.1 data, please see this collection.
You can learn more about Common Pile in our paper.
Mixing rates and token counts
The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage.
During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.telegram-news-ua-dataset
Aisberg Telegram News UA
A continuously updated, de-identified corpus of Ukrainian Telegram news and the discussion around it, published by the Ukrainian non-profit Aisberg (ГО «АЙЗБЕРГ»). It comes in two layers. The first is the raw monthly stream: every post from a fixed set of public news channels, with its reactions and its comment thread. The second is the analysis behind every report Aisberg publishes: posts from different channels grouped into one event, the manipulation… See the full description on the dataset page: https://huggingface.co/datasets/aisbergpublicorganization/telegram-news-ua-dataset.bio-mqm-datasetThis dataset is compiled from the official Amazon repository (all respective licensing applies).
It contains system translations, multiple references, and their quality evaluation on the MQM scale. It accompanies the ACL 2024 paper Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains.
Watch a brief 4 minutes-long video.
Abstract: We introduce a new, extensive multidimensional quality metrics (MQM) annotated dataset covering 11 language pairs in the biomedical domain. We use this… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/bio-mqm-dataset.FlashRAG_datasets
⚡FlashRAG: A Python Toolkit for Efficient RAG Research
FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms.
With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components.
For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.ua-open-data
Україна: дзеркало відкритих даних (data.gov.ua)
Автоматичне дзеркало публічних наборів data.gov.ua,
яке підтримує пайплайн JoTalbot/ukraine.
Набори
Набір
Файлів
Джерело
Єдиний державний реєстр юридичних осіб, фізичних осіб-підприємців та громадських формувань
6
—
Реєстр декларацій родинних зв’язків та доброчесності
14
—
Державний судновий реєстр України
9
—
Публічні закупівлі на сайті Prozorro
1
—
Інформація щодо стану розгляду справ
5
—… See the full description on the dataset page: https://huggingface.co/datasets/JoTalbot/ua-open-data.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.xfield-radar-dataset-20260915
XField radar dataset — formal snapshot, 2026-09-15
Upload status: COMPLETE — every shard verified against its remote SHA-256 and byte size. See UPLOAD_COMPLETE.json.
This public research snapshot preserves the currently admitted XField/GRT simulation dataset: radar inputs, existing GT, available raw sensor products, provenance, and the frozen N141 train/validation split. It is not a claim that historical labels meet the newly repaired independent dense-GT pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/KAS2003/xfield-radar-dataset-20260915.yayi2_pretrain_data
介绍/Introduction
本数据集源自雅意训练语料,我们精选了约100B数据,数据大小约为500GB。我们期望通过雅意预训练数据的开源推动中文预训练大模型开源社区的发展,并积极为此贡献力量。通过开源,我们与每一位合作伙伴共同构建雅意大模型生态。
We opensource the pre-trained dataset in this release, it should contain more than 100B tokens depending on the tokenizer you use, requiring more than 500GB of local storage. By open-sourcing the pre-trained dataset, we aim to contribute to the development of the Chinese pre-trained large language model open-source community. Through open-source, we aspire to… See the full description on the dataset page: https://huggingface.co/datasets/wenge-research/yayi2_pretrain_data.Dataset
MM-OphBench: Multi-Center Multimodal Clinical Ophthalmic Benchmark Dataset
A Large-Scale, Standardized Multi-Center Benchmark Covering 7 Imaging Modalities & 4.3M+ Clinical Records
1. Executive Summary & Repository Overview
The MM-OphBench repository hosts a petabyte-scale, clinically harmonized ophthalmic image archive compiled from leading ophthalmic hospitals and benchmark cohorts. It spans 4,307,415 high-resolution diagnostic images and multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Kaphathy/Dataset.dapo-math-17kInternData-fractal20220817_dataturkish-llm-dataset
Turkish Pretraining Corpus
Dataset Description
This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models.
This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/tascib/turkish-llm-dataset.RoboInter-Data
RoboInter-Data: Intermediate Representation Annotations for Robot Manipulation
Rich, dense, per-frame intermediate representation annotations for robot manipulation, built on top of DROID and RH20T. Developed as part of the RoboInter project. You can try our Online demo.
The annotations cover 230k episodes and include: subtasks,
primitive skills, segmentation, gripper/object bounding boxes, placement proposals, affordance boxes,
grasp poses, traces, contact points, etc. And each… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/RoboInter-Data.PortPy_Dataset
PortPy: Planning and Optimization for Radiation Therapy
Data Overview
PortPy equips researchers with a robust benchmark patient dataset, sourced from the FDA-approved Eclipse commercial treatment planning system through its API. This dataset embodies all necessary elements for optimizing various machine configurations such as beam angles, aperture shapes, and leaf movements. It includes
Dose Influence Matrix (AKA dose deposition matrix, dij matrix): The dose… See the full description on the dataset page: https://huggingface.co/datasets/PortPy-Project/PortPy_Dataset.minWM-dataAegis-AI-Content-Safety-Dataset-2.0
🛡️ Nemotron Content Safety Dataset V2
The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1.
To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0.Nemotron-VLM-Dataset-v2
Nemotron-VLM-Dataset v2
Versions
Date
Commit
Changes
2025-11-05
head
Fix nights_cot dataset. Fix/filter broken <think> entries. Update fintabnet instructions. Update indexes.
2025-10-28
214051e
Initial Release
Dataset Description
Following up on Llama Nemotron VLM Dataset V1 with 3 million samples, we are releasing the Nemotron VLM Dataset V2 with almost three times as many high-quality samples.
This time, our focus was on three… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-VLM-Dataset-v2.LEMAS-Dataset-train
Overview
This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project).
It contains a large-scale training set (150k+ hours) and a curated evaluation set
(500 utterances per language) covering 10 languages, all with word-level alignment.
Fields
key: unique utterance identifier; the first two characters indicate the language ID
audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-train.knowref_60k_raw
The Knowref 60K Dataset
Project: https://github.com/aemami1/KnowRef60k
Data source: https://github.com/aemami1/KnowRef60k/tree/28e5385d17967744ccb3bdba45fdd89d9690307d
Fields
annotation_strength (str): annotator agreement from 1-5
candidate_0 (str): the first candidate name
candidate_1 (str): the second candidate name
original_sentence (str): sentence before swapping the names
swapped_sentence (str): sentence after swapping the names with square brackets marking the… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/knowref_60k_raw.data_4
SpatialEncoder WDS release (in progress)
This repository contains a partition of spatialencoder-wds-native-v1, released
as uncompressed WebDataset tar shards, normally about 1 GiB. All five
repositories are parts of the same release; consult each manifest.json.
The manifest lists only uploaded shards whose remote size and SHA-256 have
been verified. An incomplete manifest is not a complete dataset.
New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds4/data_4.instruction-datasetThis is the blind eval dataset of high-quality, diverse, human-written instructions with demonstrations. We will be using this for step 3 evaluations in our RLHF pipeline.
