datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dcvlm-balanced-200b
DCVLM-Balanced (200B tokens)
DCVLM-Balanced is the balanced-mixture training set from our DataComp-VLM paper.
It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat
WebDataset tar shards so it can be consumed by any training
stack.
This is a 200B-token release consisting of 112,358,849 samples, curated from our DCVLM-large data pool.
The instruction-heavy counterpart (DCVLM-baseline) is available as
dcvlm-baseline-200b, along with… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-balanced-200b.ai2thor-perspective-qa-20k-balanced-splits-with-objai2thor-perspective-qa-100k-balanced-training-v1-splitskl3m-data-sample-005-balancedragbench-sentence-relevance-balanceddigit_mask_ensemble_distilled_from_cv12_balanced_mfcc
Dataset Card for "digit_mask_ensemble_distilled_from_cv12_balanced_mfcc"
More Information needed
procthor-100-counting-balancedvctk_resampled_16k_balancednvidia_open_reasoning_balanced_100k
nvidia_open_reasoning_balanced_100k
A domain-balanced 100k reasoning SFT dataset built from three NVIDIA Open Reasoning
datasets: 33,333 examples each for math, code, and science (99,999 total).
Each example is a single-turn conversation with a full reasoning trace:
conversations: [
{"from": "human", "value": "<problem>"},
{"from": "gpt", "value": "<think>\n<reasoning trace>\n</think><final solution>"}
]
Columns
column
description
conversations… See the full description on the dataset page: https://huggingface.co/datasets/bhheo/nvidia_open_reasoning_balanced_100k.gtex-10M-balanced-tiles
GTEx 10M Balanced Tiles
This dataset contains 10,000,000 JPEG-encoded 224x224 pathology tiles from GTEx SVS slides in s3://path-datasets/gtex/svs_by_tissue, balanced at 250,000 tiles for each of the 40 tissue prefixes. Source objects under 80,000,000 bytes are ignored because the GTEx prefix contains tiny unsupported SVS objects. Tiles are sampled from virtual OpenSlide levels 0, 1, and 2, so the effective field of view varies while the emitted JPEG size stays fixed.
Each… See the full description on the dataset page: https://huggingface.co/datasets/medarc/gtex-10M-balanced-tiles.0_digit_mask_ensemble_distilled_from_cv12_balanced_mfcc
Dataset Card for "0_digit_mask_ensemble_distilled_from_cv12_balanced_mfcc"
More Information needed
everyayah_curated_1s_20s_balancedrl__24GPU_base__mix_h2_language_balanced__r2egym-nl2bash-stackabir177m-pretrain-balanced20-ezhijaru
abir177m pretrain mix — balanced20 en/zh/hi/ja/ru
Frozen packed-token shards for reproducible abir177m GPT-2–style pretraining.
Languages: 20% each en, zh, hi, ja, ru
Source streams: FineWeb (en) + FineWeb-2 (zh/hi/ja/ru)
Tokenizer: mistralai/Mistral-Nemo-Base-2407
Packing: 2048-token causal LM blocks (input_ids, labels identical)
Target budget: 3.55B tokens (1,733k sequences)
See meta.json for exact mixture + dataset map + seed.
ai2thor-perspective-qa-balanced-400-v2ai2thor-perspective-qa-800-balanced-val-v11_digit_mask_ensemble_distilled_from_cv12_balanced_mfcc
Dataset Card for "1_digit_mask_ensemble_distilled_from_cv12_balanced_mfcc"
More Information needed
everyayah_curated_1s_20s_balanced_largeai2thor-perspective-qa-400-balanced-diverseai2thor-perspective-qa-400-balanced-v2refchartqa-balanced-10kai2thor-perspective-qa-800-balanced-diverse-v1ai2thor-perspective-qa-400-balanced-diverse-v2AudioSet-Strong-Balancedtwist_subset_balanced_100k_448_multi_repo_viewerfix_rg50This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 300,
"total_frames": 100000,
"total_tasks": 33245,
"chunks_size": 1000,
"data_files_size_in_mb": 300,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:300"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/lyl472324464/twist_subset_balanced_100k_448_multi_repo_viewerfix_rg50.airbus-balanced-subsetnemotron-sft-balanced-2b-v1
Nemotron SFT Dataset
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Statistics
Total Samples: 200,000
Total Tokens: 1,252,287,904
Average Tokens per Sample: 6261.4
Tokenizer: Qwen/Qwen3-0.6B
Random Seed: 42
Strategy: balanced
Subset Distribution
Subset
Samples
Tokens
Target
Completion
Avg Tokens/Sample
Stage-1/math
20,000
151,546,125
20,000… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-balanced-2b-v1.BALANCED_MSPP_MSPI_IEMOroute_subtasks_dual_overhead_pi05_448_dagger_interventions_balanced
Route Subtasks + Balanced DAgger Interventions, 448px
This is a new LeRobot v2.1 dataset derived from
ajaysri/route_subtasks_dual_overhead_pi05_448 and a subsequent
DAgger collection. The original dataset is not modified.
The base contributes 495 episodes and 128,742 frames. Only
frames recorded while the human collector was actively intervening are added;
autonomous policy-control frames and policy_target_action are excluded from
the training targets. The DAgger action column… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/route_subtasks_dual_overhead_pi05_448_dagger_interventions_balanced.xray-balanced-dataset
