datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dclm-baseline-1.0
DCLM-baseline
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets
Llama2
7B
2T
✗
49.2
45.8
34.1
DeepSeek
7B
2T
✗
50.7
48.5
35.3
Mistral-0.3
7B
?
✗
57.0
62.7
45.1
QWEN-2
7B
?
✗
57.5
71.9
50.5
Llama3
8B
15T
✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.dclm-baseline-1.0-parquet
DCLM-baseline
Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format.
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.droid_1.0.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/droid_1.0.1.magpie-ultra-v1.0
Dataset Card for magpie-ultra-v1.0
This dataset has been created with distilabel.
Dataset Summary
magpie-ultra it's a synthetically generated dataset for supervised fine-tuning using the Llama 3.1 405B-Instruct model, together with other Llama models like Llama-Guard-3-8B and Llama-3.1-8B-Instruct.
The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing… See the full description on the dataset page: https://huggingface.co/datasets/argilla/magpie-ultra-v1.0.libero_spatial_no_noops_1.0.0_lerobotlibero_object_no_noops_1.0.0_lerobotlibero_10_no_noops_1.0.0_lerobotlibero_goal_no_noops_1.0.0_lerobotdroid_1.0.1_v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95584,
"total_frames": 27607757,
"total_tasks": 49596,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95584"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30.dAgger_build_block_tower_1.0.0-advantages
Advantage Values for villekuosmanen/dAgger_build_block_tower_1.0.0
Pre-computed advantage values for offline RL training.
Source
Dataset: villekuosmanen/dAgger_build_block_tower_1.0.0
Value Model: villekuosmanen/rewact_build_block_tower_all_3
N-step lookahead: 50
Files
This dataset contains per-episode parquet files with advantage values for each frame.
Usage
from pathlib import Path
import pandas as pd
# Load advantages for a specific episode… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/dAgger_build_block_tower_1.0.0-advantages.droid_1.0.1_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/droid_1.0.1_test.droid_1.0.1_v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/muacha/droid_1.0.1_v30.chat1.0phail-v1.0
PhAIL: Real-Robot VLA Evaluation Benchmark (v1.0)
This dataset accompanies an anonymous submission to the NeurIPS 2026
Evaluations and Datasets track. The paper, code, and dataset are all under
double-blind review; identifying URLs and author information have been
withheld.
PhAIL is a real-robot evaluation benchmark for vision-language-action (VLA)
policies. It contains synchronized exterior and wrist RGB video, end-effector
and gripper telemetry, and per-rollout event… See the full description on the dataset page: https://huggingface.co/datasets/phail-anon/phail-v1.0.Ornith-1.0-9B-atlas
juiceb0xc0de/Ornith-1.0-9B-atlas
A brain atlas for deepreinforce-ai/Ornith-1.0-9B, the 9B agentic-coding model that reports SOTA results on Terminal-Bench, SWE-Bench, and other agentic coding benchmarks. This is not a chat dataset or a benchmark — it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know why this model survives surgical edits, where… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Ornith-1.0-9B-atlas.droid_1.0.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/droid_1.0.1.Pluto-Nano-1.0-Pretrain-v2
ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2)
Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI).
v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.droid_1.0.1lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private
Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v1.0
Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v1.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private.droid_1.0.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95617,
"total_frames": 27618651,
"total_tasks": 49611,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95617"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ygtxr1997/droid_1.0.1.droid_1.0.1
DROID 1.0.1 for pi05 training: successful, non-idle segments
lerobot/droid_1.0.1 rewritten twice, videos untouched. First
augment_droid_delta_ee_gripper_events.py (lerobot_policy_framepick) added action.delta_ee,
observation.extrinsics.static1/static2/wrist1 and the observation.gripper.time_* columns;
the flat action is the delta-EE command and the flat observation.state the EE state. Then
benchmarks/robolab/augment_droid_dataset.py (this revision) mirrored the data pipeline of… See the full description on the dataset page: https://huggingface.co/datasets/dgrachev/droid_1.0.1.dclm-baseline-1.0_subset_30MPHI-SPIKE-C172x-Community-Dataset-v1.0
PHI-SPIKE C172X Community Dataset v1.0
Dataset Summary
PHI-SPIKE C172X Community Dataset v1.0 is a simulation-based aerospace Prognostics and Health Management (PHM) dataset and training-artifact release developed from the PHI-SPIKE C172X research campaign.
The release provides:
JSBSim C172X reference telemetry;
benchmark metadata;
training histories;
trained PyTorch model checkpoints;
per-run evaluation metrics; and
five-seed campaign summaries.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/PHI-SPIKE-C172x-Community-Dataset-v1.0.4B-predict.rule-r-1.0-k-256.L-1024.statml-arxivdroid_1.0.1_v30_merged
Leyo/droid_1.0.1_v30_merged
In-place merged per chunk.
One row per episode in each data/chunk-XXX/file-YYY.parquet.
Timestamp-level columns are lists; episode-level fields are meta__*.
Join key: episode_index.
Generated with Polars (lazy/streaming), parallelized across files.
Parquet compression: zstd.
3_4_fusechat_v1_openchat-3.5_mixtral-8x7b-instruct-v0.1_solar-10.7b-instruct-v1.0_representationFnii-VLA-Kinova-1.0droid_1.0.1_v30_compact_5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30_compact_5.Ko-Agent-Trajectories-1.0
Ko-Agent-Trajectories-1.0
Dataset card v1.1.1 (2026-09-22). The pipeline code is now released in this repository
under pipeline/, together with the API catalogue, the scenario templates and the complete
prompt set. The card reports the completed human review study and the v1.1 artefacts
(behaviour DPO config, per-item validation scores, manifest, filter asset).
Korean edition: README.ko.md.
TL;DR
A Korean multi-turn agent ↔ tool trajectory corpus synthesized… See the full description on the dataset page: https://huggingface.co/datasets/taejoon89/Ko-Agent-Trajectories-1.0.Qiita-1.07MThis dataset contains 1,074,174 articles published on Qiita.
