datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dclm-baseline-1.0
DCLM-baseline
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets
Llama2
7B
2T
✗
49.2
45.8
34.1
DeepSeek
7B
2T
✗
50.7
48.5
35.3
Mistral-0.3
7B
?
✗
57.0
62.7
45.1
QWEN-2
7B
?
✗
57.5
71.9
50.5
Llama3
8B
15T
✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.droid_1.0.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Franka",
"total_episodes": 95600,
"total_frames": 27612581,
"total_tasks": 0,
"total_videos": 286800,
"total_chunks": 95,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:95600"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1.essential-web-v1.0
🌐 Essential-Web: Complete 24-Trillion Token Dataset
🏆 Website | 🖥️ Code | 📖 Paper | ☁️ AWS
📋 Dataset Description
Essential-Web is a 24-trillion-token web dataset with document-level metadata designed for flexible dataset curation. The dataset provides metadata including subject matter classification, web page type, content complexity, and document quality scores for each of the 23.6 billion documents.
Researchers can filter and curate specialized datasets… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/essential-web-v1.0.dclm-baseline-1.0-parquet
DCLM-baseline
Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format.
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.droid_1.0.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/droid_1.0.1.sdxl-1.0
check sdxl.parrotzone.art for easy viewing ⋆。°✩
all images were made with SDXL 1.0 + the 0.9 VAE
steps: 20
cfg scale: 7
no refiner
random seeds
magpie-ultra-v1.0
Dataset Card for magpie-ultra-v1.0
This dataset has been created with distilabel.
Dataset Summary
magpie-ultra it's a synthetically generated dataset for supervised fine-tuning using the Llama 3.1 405B-Instruct model, together with other Llama models like Llama-Guard-3-8B and Llama-3.1-8B-Instruct.
The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing… See the full description on the dataset page: https://huggingface.co/datasets/argilla/magpie-ultra-v1.0.v1.0
MuLAn: : A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation
MuLAn is a novel dataset comprising over 44K MUlti-Layer ANnotations of RGB images as multilayer, instance-wise RGBA decompositions, and over 100K instance images. It is composed of MuLAn-COCO and MuLAn-LAION sub-datasets, which contain a variety of image decompositions in terms of style, composition and complexity. With MuLAn, we provide the first photorealistic resource providing instance… See the full description on the dataset page: https://huggingface.co/datasets/mulan-dataset/v1.0.LET-KUAVO-VLA-1.0-Dataset
LET-KUAVO-VLA-1.0-Dataset
libero_spatial_no_noops_1.0.0_lerobotlibero_object_no_noops_1.0.0_lerobotlibero_10_no_noops_1.0.0_lerobotlibero_goal_no_noops_1.0.0_lerobotNemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.RealCADBench-V1.0
RealCADBench-V1.0
RealCADBench-V1.0 contains design-intent inputs and ground-truth STL shapes for evaluating AI CAD models and agents on parts and assemblies.
The test split is an evaluation collection, not a newly created train/test partition. The default all configuration combines every sample. The six other configurations provide the individual subsets without duplicating the Parquet files.
Category
Subset
Samples
part
text
568
part
2d_drawing
236
part
real_pic… See the full description on the dataset page: https://huggingface.co/datasets/RealCADBench/RealCADBench-V1.0.Aegis-AI-Content-Safety-Dataset-1.0
🛡️ Nemotron Content Safety Dataset V1
Nemotron Content Safety Dataset V1, formerly known as Aegis AI Content Safety Dataset, is an open-source content safety dataset (CC-BY-4.0), which adheres to Nvidia's content safety taxonomy, covering 13 critical risk categories (see Dataset Description).
Dataset Details
Dataset Description
Nemotron Content Safety Dataset V1 is comprised of approximately 11,000 manually annotated interactions between humans and LLMs, split… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-1.0.occiglot-fineweb-v1.0
Occiglot Fineweb v1.0
We present a more mature version of the multilingual Occiglot Fineweb corpus. In this early form, the dataset contains roughly 430M heavily cleaned documents from 10 languages.
Occiglot Fineweb builds on our existing collection of curated datasets and pre-filtered web data.
Subsequently, all documents were filtered with language-specific derivatives of the fine-web processing pipeline and different levels of depuplicated.
We provide the data at 3 levels of… See the full description on the dataset page: https://huggingface.co/datasets/occiglot/occiglot-fineweb-v1.0.DOTAv1.0vlm_evaluation_v1.0
Datacard
This dataset is the evaluation VLM dataset used in VLABench. It is designed to evaluate the planning capabilities of Vision-Language Models (VLMs) in embodied scenarios.
Source
Project Page: https://vlabench.github.io/
Arxiv Paper: https://arxiv.org/abs/2412.18194
Code: https://github.com/OpenMOSS/VLABench
Uses
The dataset structure is as follows:
vlm_evaluation_v1.0/
├── CommenSence/
├── add_condiment_common_sense/
├──… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlm_evaluation_v1.0.peoples_speech_v1.0
Dataset Card for People's Speech
Dataset Summary
The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech_v1.0.droid_1.0.1_v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95584,
"total_frames": 27607757,
"total_tasks": 49596,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95584"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30.dAgger_build_block_tower_1.0.0-advantages
Advantage Values for villekuosmanen/dAgger_build_block_tower_1.0.0
Pre-computed advantage values for offline RL training.
Source
Dataset: villekuosmanen/dAgger_build_block_tower_1.0.0
Value Model: villekuosmanen/rewact_build_block_tower_all_3
N-step lookahead: 50
Files
This dataset contains per-episode parquet files with advantage values for each frame.
Usage
from pathlib import Path
import pandas as pd
# Load advantages for a specific episode… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/dAgger_build_block_tower_1.0.0-advantages.droid_1.0.1_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/droid_1.0.1_test.spandan-1M-V1.0-raw
Spandan
A Large Photoplethysmography (PPG) Signal Dataset of 1 Million+ Indian Subjects
In Sanskrit, "Spandan" (स्पन्दन - spandana) represents one of the most fundamental aspects of existence - the rhythmic pulsation that permeates all life. Derived from the root verb "spand" (स्पन्द), meaning "to throb" or "to pulsate," it beautifully captures the essence of the heartbeat.
Dataset Overview
Spandan is an extensive repository containing over 1 million… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/spandan-1M-V1.0-raw.bird-critic-1.0-sqlite
📢 Update 2026-03-23
We release BIRD-Critic-SQLite, a dataset containing 500 high-quality user issues focused on real-world SQLite database applications. Along with the dataset, we also release three RL-trained models: BIRD-Talon-14B, BIRD-Talon-7B, and BIRD-Zeno-7B. The schema file is included in the code repository https://github.com/bird-bench/BIRD-CRITIC-1/blob/main/baseline/data/sqlite_schema.jsonl
BIRD-CRITIC-1.0-SQLite
BIRD-Critic is the first SQL debugging… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird-critic-1.0-sqlite.AAAR-1.0 AAAR-1.0 Benchmark
📄 Paper: https://huggingface.co/papers/2410.22394
🌐 Website: https://renzelou.github.io/AAAR-1.0/
🤗 This repository contains the AAAR-1.0 benchmark dataset.
🚨 Please DO NOT use the data for training!
Get Performance on AAAR-1.0
Please refer to our code repository for detailed instructions on running various LLMs on the AAAR-1.0 benchmark, and report the performances.
Data Details
1. Equation Inference 🌟:… See the full description on the dataset page: https://huggingface.co/datasets/Reza8848/AAAR-1.0.cv-v1.0-segment
CommonVoice v1 Phone-Segment Alignments
Phone-level time alignments for 10 languages of Mozilla Common Voice,
packaged in a canonical segmentation schema with embedded 16 kHz audio. The
phone boundaries come from the charsiu/cv_ali
release of MFA alignments; the audio and transcripts come from
Common Voice Corpus 13.0 (2023-03-09).
Dataset summary
lang
train rows
train hrs
val rows
val hrs
test rows
test hrs
en
1,008,669
1,354.0
3,537
4.9
1,285
1.7
rw… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/cv-v1.0-segment.more_peoples_speech_v1.0
Dataset Card for People's Speech
Dataset Summary
The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/chris745820/more_peoples_speech_v1.0.hle_labeled-v1.0
HLE Labeled Dataset
このデータセットは「Humanity’s Last Exam」ベンチマーク用のデータセットの
categoryにsubcategoryを追加したものです。subcategoryの分類ラベルはqwen/qwen3-235b-a22bで生成しています。
モデルのカテゴリ別の評価に利用するのが目的です。
データ構造
変更点はもともとのHLEにsubcategoryフィールド追加したのみです。
id: レコードのユニークID
question: 問題文(文字列)
answer: 正解
answer_type: "exactMatch" などの解答形式
rationale: 解答手順・根拠
category: 大分類(例: "Math")
subcategory: 小分類のリスト(例: ["Math/Number Theory","Math/Discrete Mathematics"])
image, image_preview, rationale_image:… See the full description on the dataset page: https://huggingface.co/datasets/LLMcompe-Team-Watanabe/hle_labeled-v1.0.dclm-baseline-1.0-llama3-tokenized-shuffled
!! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !!
DCLM-Baseline Pretokenized (LLaMA 3.1, 8192 context)
This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines.
The original DCLM-Baseline… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled.
