datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vchitect_T2V_DataVerse
Vchitect-T2V-Dataverse
Vchitect Team1
1Shanghai Artificial Intelligence Laboratory
Paper |
Project Page |
Data Overview
The Vchitect-T2V-Dataverse is the core dataset used to train our text-to-video diffusion model, Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.
It comprises 14 million high-quality videos collected from the Internet, each paired with detailed textual… See the full description on the dataset page: https://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse.Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.danbooru
Danbooru 2024 Dataset
Danbooru 2024 数据集
A collection of images from Danbooru website, organized and packaged by ID sequence. This dataset is for research and learning purposes only.
本数据集收集了来自 Danbooru 网站的图像,按 ID 顺序组织打包。该数据集仅用于研究和学习目的。
Dataset Description
数据集描述
This dataset contains image resources from Danbooru website, updated to ID 8380648 (Update time: 2024-11-03).
本数据集包含来自 Danbooru 网站的图像资源,更新至 ID 8380648(更新时间:2024-11-03)。
Data… See the full description on the dataset page: https://huggingface.co/datasets/picollect/danbooru.FTP-1-Dataset
FTP-1-Dataset
FTP-1-Dataset contains heterogeneous tactile manipulation data for FTP-1 pretraining.
This release currently includes 18 dataset archives:
FreeTacMan
MotionTrans
RDP
RDP_Bimanual
RH20TCfg5Franka
RH20TCfg6ATIAxia
RH20TCfg7Tactile
Unit
Unit_Bimanual
VLA_touch
ViTaMIn
VisuoTactile_D-WHEEL
VisuoTactile_QINGLOONG
exUMI
sharpa
Each dataset directory contains either a single <dataset>.tar file or split parts named <dataset>.tar.part-*. For split archives, concatenate… See the full description on the dataset page: https://huggingface.co/datasets/MJJJJ1064/FTP-1-Dataset.X-Voice-Dataset-Train
X-Voice Training Dataset
Overview
The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling.
Also the train set of X-Voice Model.
Core Statistics
Total Speech Duration: 420K hours
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.robopoint-data
RoboPoint Dataset Card
Dataset details
This dataset contains 1432K image-QA instances used to fine-tune RoboPoint, a VLM for spatial affordance prediction. It consists of the following parts:
347K object reference instances from a synthetic data pipeline;
320K free space reference instances from a synthetic data pipeline;
100K object detection instaces from LVIS;
150K GPT-generated instruction-following instances from liuhaotian/LLaVA-Instruct-150K;
515K general-purpose… See the full description on the dataset page: https://huggingface.co/datasets/wentao-yuan/robopoint-data.ReCo-Data
ReCo-Data Dataset Card
Introduction
ReCo-Data is a large-scale, high-quality video editing dataset comprising 500K+ instruction-video pairs. This card provides its statistics, collection pipeline, and dataset format.
1. Dataset Statistics
Statistics
Figure Caption:
(a) Overview of scale
(b) Task distribution showing balanced quantities: Replace (156.6K), Style (130.6K), Remove (121.6K), and Add (115.6K). Human evaluation on 200 randomly… See the full description on the dataset page: https://huggingface.co/datasets/HiDream-ai/ReCo-Data.medpmc-11m-dataset_jun24_baseline
MedPMC WebDataset
MedPMC is a large-scale medical image-text dataset curated from articles in the PubMed Central (PMC) collection. This release contains approximately 11 million image-text pairs collected from the June 2024 PMC baseline. MedPMC is an ongoing effort, and future releases will continue to expand the dataset with newly published literature, improved annotations, and additional resources.
This dataset is presented in the paper MedPMC: A Systematic Framework for… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-11m-dataset_jun24_baseline.olmoearth_pretrain_datasetThis is the pre-training dataset for training the OlmoEarth pre-trained remote sensing foundation models.
Documentation is on GitHub at https://github.com/allenai/olmoearth_pretrain/blob/main/docs/Pretraining-Dataset.md
The dataset is released under CC BY 4.0. It includes data from the following sources:
Sentinel-2 L2A imagery from the European Space Agency, available under the Copernicus Sentinel Data and Service Legal Notice
Sentinel-1 GRD IW vv+vh imagery from the European Space Agency… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmoearth_pretrain_dataset.findtextCenterNet_datasetdanbooru2024-webp-4Mpixel
🎨 Danbooru2024 Webp 4MPixel Dataset
📊 Dataset Overview
The Danbooru2024-Webp dataset is a comprehensive collection focused on animation and illustration artwork, derived from the official Danbooru platform. It contains approximately 8.05 million high-quality, user-annotated images with corresponding tags and textual descriptions.
This dataset is 4MP-focused webp resized-dataset of Danbooru2024.
✨ Features
📋 Metadata Support
Includes a Parquet… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/danbooru2024-webp-4Mpixel.X2Edit-Dataset
X2Edit
Introduction
X2Edit Dataset is a comprehensive image editing dataset that covers 14 diverse editing tasks and exhibits substantial advantages over existing open-source datasets including AnyEdit, HQ-Edit, UltraEdit, SEED-Data-Edit, ImgEdit and OmniEdit.
For the relevant data construction scripts, model training and inference scripts, please refer to X2Edit.
News
2025/09/16: We are about to release a dataset constructed by Qwen-Image and… See the full description on the dataset page: https://huggingface.co/datasets/OPPOer/X2Edit-Dataset.scientific-stuff-1raw_primitive_datasets
Datacard
This is the official fine-tuning dataset provided by VLABench (raw data), with 500 episodes each task. The current version includes 10 primitive tasks.
Source
Project Page: https://vlabench.github.io/
Arxiv Paper: https://arxiv.org/abs/2412.18194
Code: https://github.com/OpenMOSS/VLABench
Uses
Download all archive files and use the following command to extract:
cat vlabench_primitive.tar.gz.* | tar -xzvf -
In the resulting… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/raw_primitive_datasets.describe-anything-dataset
Describe Anything: Detailed Localized Image and Video Captioning
NVIDIA, UC Berkeley, UCSF
Long Lian, Yifan Ding, Yunhao Ge, Sifei Liu, Hanzi Mao, Boyi Li, Marco Pavone, Ming-Yu Liu, Trevor Darrell, Adam Yala, Yin Cui
[Paper] | [Code] | [Project Page] | [Video] | [HuggingFace Demo] | [Model/Benchmark/Datasets] | [Citation]
Dataset Card for Describe Anything Datasets
Datasets used in the training of describe anything models (DAM).
The datasets are in tar files. These… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/describe-anything-dataset.danbooru2023-webp-4Mpixel-224The data set is just resized to 224*224
https://huggingface.co/datasets/KBlueLeaf/danbooru2023-webp-4Mpixel
Pseudo code for processing
def resize_image(file_path):
with Image.open(file_path) as img:
resized_img = img.resize((224, 224))
resized_img.save(file_path)
voice-data
Voice-Data: a curated multi-corpus voice dataset for voice–text contrastive training
voice-data is a single, globally-shuffled WebDataset that bundles several voice/speech corpora into one ready-to-train mixture for voice–text contrastive (CLAP-style) models such as VoiceCLAP. Each clip pairs 48 kHz mono FLAC audio with a natural-language text caption describing the voice — its emotion, prosody, timbre, speaking style, recording context, and speaker traits.
The distinguishing… See the full description on the dataset page: https://huggingface.co/datasets/gijs/voice-data.Ola-DataThis repository contains the data presented in Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment.
Code: https://github.com/Ola-Omni/Ola
majestrino-dataimagenet-1k-wds
Dataset Summary
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated.
💡… See the full description on the dataset page: https://huggingface.co/datasets/dark-xet/imagenet-1k-wds.STRIDE-QA-Dataset
STRIDE-QA Dataset
📦 Dataset
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
Category
Description
Object-centric Spatial QA
Spatial relations between two… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset.dan-webp-newgptsovits_dataset
bhyuan/gptsovits_dataset
GPT-SoVITS speech dataset, packed as WebDataset tar shards.
Layout
data/
train/
metadata.csv
audio/
train-000.tar
train-001.tar
...
validation/
metadata.csv
audio/
validation-000.tar
...
test/
metadata.csv
audio/
test-000.tar
...
Shard counts:
youshengshu_v5_test: 6536 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/bhyuan/gptsovits_dataset.commonvoice22-sidon-dacvae
CommonVoice 22 (Sidon-enhanced) converted to DAC VAE latents
Source
sarulab-speech/commonvoice22_sidon
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/commonvoice22-sidon-dacvae.DataCompDR-12M-bf16
Dataset Card for DataCompDR-12M-BFloat16
This dataset contains synthetic captions, embeddings, and metadata for DataCompDR-12M.
The metadata has been generated using pretrained image-text models on a 12M subset of DataComp-1B.
For details on how to use the metadata, please visit our github repository.
The dataset with the original captions is now available at mlfoundations/DataComp-12M.
The UIDs per shards match between mlfoundations/DataComp-12M and apple/DataCompDR-12M-bf16.… See the full description on the dataset page: https://huggingface.co/datasets/apple/DataCompDR-12M-bf16.pure-danbooru
pure-danbooru Dataset
WarriorMama777/pure-danbooru
概要
このデータセットはDanbooruをメインに700万枚のイラストで構成された大規模な画像データセットです。
特徴は以下の通りです。
Danbooruのタグを純粋に維持した未検閲のトレーニング用データセット。Danbooruのユーザーたちが地道に長年作業してきたタグ付けの作業がピュアに維持されています。
丁寧な前処理、およびトレーニング用にデータを整形済み。sd-scriptsのFinetuningに準拠した形でトレーニング用Jsonが提供されており、箱から出してすぐに使えます。
データセットの詳細
画像は主にDanbooruのイラストレーションで構成されています(Danbooruの一部のタグは隠されており有料会員しか閲覧できないので、その部分をGelbooru等から補填しています)。
画像の収集… See the full description on the dataset page: https://huggingface.co/datasets/WarriorMama777/pure-danbooru.Harmonizer-Dataset
HARMONIZER DATASET
Dataset Description
Training dataset for DiffusionHarmonizer: a generative AI model for image and video enhancement bridging neural reconstruction and photorealistic simulation .
Model checkpoints: https://huggingface.co/nvidia/Harmonizer/Training code: https://github.com/NVIDIA/harmonizer/
The dataset was curated to support the following functions of the model:
3D reconstruction artifact removal
Harmonization of inserted objects to blend… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Harmonizer-Dataset.tinyllama_pretrained_dataRaCig-DataJANuS_datasetThis repository hosts the JANuS (Joint Annotations and Names) dataset introduced in the 2023 paper Distributionally Robust Classification on a Data Budget.
As of this writing, ours is the only public dataset which is both fully annotated with ground-truth labels and fully captioned with web-scraped captions.
It is designed to be used for controlled experiments with vision-language models.
What is in JANuS?
JANuS provides metadata and image links for four new training datasets; all… See the full description on the dataset page: https://huggingface.co/datasets/penfever/JANuS_dataset.
