datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenThoughts-1k-sample
[!NOTE]
We have released a paper for OpenThoughts! See our paper here.
Open-Thoughts-1k-sample
This is a 1k sample of the OpenThoughts-114k dataset.
Open synthetic reasoning dataset with high-quality examples covering math, science, code, and puzzles!
Inspect the content with rich formatting with Curator Viewer.
Available Subsets
default subset containing ready-to-train data used to finetune the OpenThinker-7B and OpenThinker-32B models:
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample.LLaVA-OneVision-1.5-Mid-Training-85M
🚀 LLaVA-One-Vision-1.5-Mid-Training-85M Dataset is being uploaded 🚀
Upload Status
All Completed: ImageNet-21k、LAIONCN、DataComp-1B、Zero250M、COYO700M、SA-1B、MINT、Obelics
📜 Cite
If you find LLaVA-One-Vision-1.5-Mid-Training-85M useful in your research, please consider to cite the following related papers:
@misc{an2025llavaonevision15fullyopenframework,
title={LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training}… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Mid-Training-85M.MINT-1T-HTML
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.dclm-baseline-1.0
DCLM-baseline
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets
Llama2
7B
2T
✗
49.2
45.8
34.1
DeepSeek
7B
2T
✗
50.7
48.5
35.3
Mistral-0.3
7B
?
✗
57.0
62.7
45.1
QWEN-2
7B
?
✗
57.5
71.9
50.5
Llama3
8B
15T
✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.preprocessed_commoncatalog-cc-byI also seperately provide just the prompts in prompts.json
keys are the image_id, and the values are the captions generated
Captions generated by moondream: vikhyatk/moondream2
Latents generated by SDXL VAE: madebyollin/sdxl-vae-fp16-fix
Embeddings generated by SigLIP: hf-hub:timm/ViT-SO400M-14-SigLIP-384
Original dataset: common-canvas/commoncatalog-cc-by
Latents f32 and embeddings are f16 bytes
Compute cost: 16x3090 for 3 day. Approximately.
10Kh-RealOmin-OpenData
Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry.
Update Notes:Stage 3 data upload completed.
13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms
Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use
3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.geonamebase_1ps2_hf1pretraining_v1-omega_booksSAGE-10k
SAGE-10k
SAGE-10k is a large-scale interactive indoor scene dataset featuring realistic layouts, generated by the agentic-driven pipeline introduced in "SAGE: Scalable Agentic 3D Scene Generation for Embodied AI". The dataset contains 10,000 diverse scenes spanning 50 room types and styles, along with 565K uniquely generated 3D objects.
🔑 Key Features
SAGE-10k integrates a wide variety of scenes, and particularly, preserves small items… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SAGE-10k.droid_1.0.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Franka",
"total_episodes": 95600,
"total_frames": 27612581,
"total_tasks": 0,
"total_videos": 286800,
"total_chunks": 95,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:95600"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1.MXStuffFastUMI_100k_lerobot
FastUMI-100K: Advancing Data-Driven Robotic Manipulation with a Large-Scale UMI-Style Dataset
[paper] [dataset]
## Overview
FastUMI-100K is a large-scale, high-quality UMI-style dataset designed for data-driven robotic manipulation learning. Featuring over **100K+ demonstration trajectories** across **54 diverse tasks** and hundreds of object types, the dataset provides multi-view wrist-mounted fisheye images and high-frequency end-effector states. To… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/FastUMI_100k_lerobot.ngii-map-full-light
ngii-map-full-light
Light point/line extract from NGII 1/1000 topographic data for Korea.
Not for shipping into GitHub — use this Hugging Face dataset instead.
CRS
Korea_2000_Central_Belt_2010 projected meters [x, y]
Layers (per region under by_region/<region>/)
Layer
Description
C023
poles (전주/통신주)
C022
lights (가로등·보안등)
A002
roads (도로 중심선)
B001_tiny
building footprints <25 m² as centroids
B002
lines (구분/재질 라인)
Also:… See the full description on the dataset page: https://huggingface.co/datasets/SKPark1/ngii-map-full-light.RekaDaily-10k-raw
RekaDaily-10k (raw)
Raw, unscripted, first-person daily-life video, collected through
Claru, Reka's data collection marketplace — recorded by
paid collectors in their own homes and workplaces on head-mounted and handheld
phones, across multiple regions.
Videos are delivered as recorded — no cuts, no trimming, no editing, no
filtering beyond basic integrity checks. A processed tier (short clips with
machine captions) is released separately under the same RekaDaily-10k prefix.… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.ABC-130k
ABC-130k
ABC-130k is the largest open-source robot teleoperation dataset. It contains
bimanual manipulation trajectories collected on two-arm YAM stations. Episodes
are distributed as MCAP files, with subtask annotations kept as separate
artifacts so they can be revised or extended independently of the underlying
episode data. For details on the accompanying paper, see abc.bot.
Please see the GitHub repo here for code to
train and deploy with this dataset.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/XDOF/ABC-130k.cifar10
Dataset Card for CIFAR-10
Dataset Summary
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images.
The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain… See the full description on the dataset page: https://huggingface.co/datasets/uoft-cs/cifar10.Leopard-Instruct
Leopard-Instruct
Paper | Github | Models-LLaVA | Models-Idefics2
Summaries
Leopard-Instruct is a large instruction-tuning dataset, comprising 925K instances, with 739K specifically designed for text-rich, multiimage scenarios. It's been used to train Leopard-LLaVA [checkpoint] and Leopard-Idefics2 [checkpoint].
Loading dataset
to load the dataset without automatically downloading and process the images (Please run the following codes with datasets==2.18.0)… See the full description on the dataset page: https://huggingface.co/datasets/wyu1/Leopard-Instruct.common_voice_17_0OpenR1-Math-220k
OpenR1-Math-220k
Dataset description
OpenR1-Math-220k is a large-scale dataset for mathematical reasoning. It consists of 220k math problems with two to four reasoning traces generated by DeepSeek R1 for problems from NuminaMath 1.5.
The traces were verified using Math Verify for most samples and Llama-3.3-70B-Instruct as a judge for 12% of the samples, and each problem contains at least one reasoning trace with a correct answer.
The dataset consists of two… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/OpenR1-Math-220k.lasa1m-annotate-part-12cad-1000-hours
CAD-1K Open v2 - 1,018.1229 Hours
509 end-to-end, single-display Windows CAD task recordings across seven CAD software families.
Each task contains:
task_desc.json - task prompt, application, reference-input paths, and expected deliverables
input_files/ - reference inputs named input.ext or input_N.ext
output_files/ - submitted CAD deliverables and supplemental outputs named output.ext or output_N.ext
rubrics.json - task-specific evaluation criteria
task_overview.pdf - review… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/cad-1000-hours.weather
Weather dataset
This public dataset contains raw weather files uploaded from the local weather/ directory.
Contents
Total files: 87,013
Total size: 278,696,918,076 bytes (about 278.7 GB)
Largest file: about 110 MB
Original local archive SHA-256: 88cdf4503fcfd3665dfb76aa324aa019c8211b534df99273911c5dc16b142eb2
Top-level groups under weather/:
radar/: 45,529 files
grib/: 39,932 files
lightning/: 1,473 files
awos_full/: 39 files
metar/: 39 files
airport_info.xlsx:… See the full description on the dataset page: https://huggingface.co/datasets/1yunyi/weather.Egocentric-100K
Egocentric-100K is the largest dataset of manual labor. You can visualize the dataset here.
Egocentric-100K is state-of-the-art in hand visibility and active manipulation density compared to previous in-the-wild egocentric datasets. The complete 30,000 frame evaluation set is available at Egocentric-100K-Evaluation.
Dataset Statistics
Attribute
Value
Total Hours
100,405
Total Frames
10.8 billion
Video Clips
2,010,759
Median Clip Length
180.0 seconds
Mean… See the full description on the dataset page: https://huggingface.co/datasets/builddotai/Egocentric-100K.2025-challenge-demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/behavior-1k/2025-challenge-demos.content-20260525cff1multilingual-speech-commands-15lang
Multilingual Speech Commands Dataset (15 Languages, Augmented)
This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification.
Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang.content-202605255e13envs_1content-202607014834
