datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vchitect_T2V_DataVerse
Vchitect-T2V-Dataverse
Vchitect Team1
1Shanghai Artificial Intelligence Laboratory
Paper |
Project Page |
Data Overview
The Vchitect-T2V-Dataverse is the core dataset used to train our text-to-video diffusion model, Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.
It comprises 14 million high-quality videos collected from the Internet, each paired with detailed textual… See the full description on the dataset page: https://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse.Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.FTP-1-Dataset
FTP-1-Dataset
FTP-1-Dataset contains heterogeneous tactile manipulation data for FTP-1 pretraining.
This release currently includes 18 dataset archives:
FreeTacMan
MotionTrans
RDP
RDP_Bimanual
RH20TCfg5Franka
RH20TCfg6ATIAxia
RH20TCfg7Tactile
Unit
Unit_Bimanual
VLA_touch
ViTaMIn
VisuoTactile_D-WHEEL
VisuoTactile_QINGLOONG
exUMI
sharpa
Each dataset directory contains either a single <dataset>.tar file or split parts named <dataset>.tar.part-*. For split archives, concatenate… See the full description on the dataset page: https://huggingface.co/datasets/MJJJJ1064/FTP-1-Dataset.X-Voice-Dataset-Train
X-Voice Training Dataset
Overview
The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling.
Also the train set of X-Voice Model.
Core Statistics
Total Speech Duration: 420K hours
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.ReCo-Data
ReCo-Data Dataset Card
Introduction
ReCo-Data is a large-scale, high-quality video editing dataset comprising 500K+ instruction-video pairs. This card provides its statistics, collection pipeline, and dataset format.
1. Dataset Statistics
Statistics
Figure Caption:
(a) Overview of scale
(b) Task distribution showing balanced quantities: Replace (156.6K), Style (130.6K), Remove (121.6K), and Add (115.6K). Human evaluation on 200 randomly… See the full description on the dataset page: https://huggingface.co/datasets/HiDream-ai/ReCo-Data.X2Edit-Dataset
X2Edit
Introduction
X2Edit Dataset is a comprehensive image editing dataset that covers 14 diverse editing tasks and exhibits substantial advantages over existing open-source datasets including AnyEdit, HQ-Edit, UltraEdit, SEED-Data-Edit, ImgEdit and OmniEdit.
For the relevant data construction scripts, model training and inference scripts, please refer to X2Edit.
News
2025/09/16: We are about to release a dataset constructed by Qwen-Image and… See the full description on the dataset page: https://huggingface.co/datasets/OPPOer/X2Edit-Dataset.olmoearth_pretrain_datasetThis is the pre-training dataset for training the OlmoEarth pre-trained remote sensing foundation models.
Documentation is on GitHub at https://github.com/allenai/olmoearth_pretrain/blob/main/docs/Pretraining-Dataset.md
The dataset is released under CC BY 4.0. It includes data from the following sources:
Sentinel-2 L2A imagery from the European Space Agency, available under the Copernicus Sentinel Data and Service Legal Notice
Sentinel-1 GRD IW vv+vh imagery from the European Space Agency… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmoearth_pretrain_dataset.findtextCenterNet_datasetmedpmc-11m-dataset_jun24_baseline
MedPMC WebDataset
MedPMC is a large-scale medical image-text dataset curated from articles in the PubMed Central (PMC) collection. This release contains approximately 11 million image-text pairs collected from the June 2024 PMC baseline. MedPMC is an ongoing effort, and future releases will continue to expand the dataset with newly published literature, improved annotations, and additional resources.
This dataset is presented in the paper MedPMC: A Systematic Framework for… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-11m-dataset_jun24_baseline.describe-anything-dataset
Describe Anything: Detailed Localized Image and Video Captioning
NVIDIA, UC Berkeley, UCSF
Long Lian, Yifan Ding, Yunhao Ge, Sifei Liu, Hanzi Mao, Boyi Li, Marco Pavone, Ming-Yu Liu, Trevor Darrell, Adam Yala, Yin Cui
[Paper] | [Code] | [Project Page] | [Video] | [HuggingFace Demo] | [Model/Benchmark/Datasets] | [Citation]
Dataset Card for Describe Anything Datasets
Datasets used in the training of describe anything models (DAM).
The datasets are in tar files. These… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/describe-anything-dataset.scientific-stuff-1gptsovits_dataset
bhyuan/gptsovits_dataset
GPT-SoVITS speech dataset, packed as WebDataset tar shards.
Layout
data/
train/
metadata.csv
audio/
train-000.tar
train-001.tar
...
validation/
metadata.csv
audio/
validation-000.tar
...
test/
metadata.csv
audio/
test-000.tar
...
Shard counts:
youshengshu_v5_test: 6536 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/bhyuan/gptsovits_dataset.robopoint-data
RoboPoint Dataset Card
Dataset details
This dataset contains 1432K image-QA instances used to fine-tune RoboPoint, a VLM for spatial affordance prediction. It consists of the following parts:
347K object reference instances from a synthetic data pipeline;
320K free space reference instances from a synthetic data pipeline;
100K object detection instaces from LVIS;
150K GPT-generated instruction-following instances from liuhaotian/LLaVA-Instruct-150K;
515K general-purpose… See the full description on the dataset page: https://huggingface.co/datasets/wentao-yuan/robopoint-data.raw_primitive_datasets
Datacard
This is the official fine-tuning dataset provided by VLABench (raw data), with 500 episodes each task. The current version includes 10 primitive tasks.
Source
Project Page: https://vlabench.github.io/
Arxiv Paper: https://arxiv.org/abs/2412.18194
Code: https://github.com/OpenMOSS/VLABench
Uses
Download all archive files and use the following command to extract:
cat vlabench_primitive.tar.gz.* | tar -xzvf -
In the resulting… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/raw_primitive_datasets.majestrino-dataOla-DataThis repository contains the data presented in Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment.
Code: https://github.com/Ola-Omni/Ola
voice-data
Voice-Data: a curated multi-corpus voice dataset for voice–text contrastive training
voice-data is a single, globally-shuffled WebDataset that bundles several voice/speech corpora into one ready-to-train mixture for voice–text contrastive (CLAP-style) models such as VoiceCLAP. Each clip pairs 48 kHz mono FLAC audio with a natural-language text caption describing the voice — its emotion, prosody, timbre, speaking style, recording context, and speaker traits.
The distinguishing… See the full description on the dataset page: https://huggingface.co/datasets/gijs/voice-data.STRIDE-QA-Dataset
STRIDE-QA Dataset
📦 Dataset
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
Category
Description
Object-centric Spatial QA
Spatial relations between two… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset.RaCig-DataDataCompDR-12M-bf16
Dataset Card for DataCompDR-12M-BFloat16
This dataset contains synthetic captions, embeddings, and metadata for DataCompDR-12M.
The metadata has been generated using pretrained image-text models on a 12M subset of DataComp-1B.
For details on how to use the metadata, please visit our github repository.
The dataset with the original captions is now available at mlfoundations/DataComp-12M.
The UIDs per shards match between mlfoundations/DataComp-12M and apple/DataCompDR-12M-bf16.… See the full description on the dataset page: https://huggingface.co/datasets/apple/DataCompDR-12M-bf16.Uni-Edit-Train-Data
Uni-Edit Training Data: Uni-Edit-148k
Project Page | GitHub Repository | Paper
👀 Intro
We introduce Uni-Edit, an intelligent image editing task that serves as the first general task for Unified Multimodal Model (UMM) tuning. Unlike conventional mixed multi-task training that suffers from inherent task conflicts and requires complex multi-stage pipelines, Uni-Edit breaks this paradigm. It achieves true mutual reinforcement by improving image… See the full description on the dataset page: https://huggingface.co/datasets/Uni-Edit/Uni-Edit-Train-Data.Harmonizer-Dataset
HARMONIZER DATASET
Dataset Description
Training dataset for DiffusionHarmonizer: a generative AI model for image and video enhancement bridging neural reconstruction and photorealistic simulation .
Model checkpoints: https://huggingface.co/nvidia/Harmonizer/Training code: https://github.com/NVIDIA/harmonizer/
The dataset was curated to support the following functions of the model:
3D reconstruction artifact removal
Harmonization of inserted objects to blend… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Harmonizer-Dataset.correctvla_datatinyllama_pretrained_dataJANuS_datasetThis repository hosts the JANuS (Joint Annotations and Names) dataset introduced in the 2023 paper Distributionally Robust Classification on a Data Budget.
As of this writing, ours is the only public dataset which is both fully annotated with ground-truth labels and fully captioned with web-scraped captions.
It is designed to be used for controlled experiments with vision-language models.
What is in JANuS?
JANuS provides metadata and image links for four new training datasets; all… See the full description on the dataset page: https://huggingface.co/datasets/penfever/JANuS_dataset.Vchitect_T2V_DataVerse_256p_8fps_wdshttps://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse resampled to 256p. Intended for training https://github.com/NilanEkanayake/TiTok-Video
scientific-stuff-2IXI-Datasetsscientific-stuff-4hf_data
