datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laions_got_talent
LAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a generated dataset comprising voice acting samples that exhibit a wide range of emotions, vocal bursts, topics, and content. This dataset is a component of the BUD-E project, spearheaded by LAION with support from Intel.
Dataset Composition
The dataset includes:
Emotional Diversity: Samples portraying various emotions to facilitate research in emotional recognition and… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent.gs-videos-v2gs-images-v2GPT-Image-Edit-1.5M
GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset
📃Arxiv | 🌐 Project Page | 💻Github
GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1.
📣 News
[2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download.
[2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.gigaspeech2
Dataset Card for GigaSpeech 2
Dataset Description
GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese.
Repository: https://github.com/SpeechColab/GigaSpeech2
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2.glint360k-wds-gz
Glint360K
This dataset is introduced in the Partial FC paper https://arxiv.org/abs/2010.05222.
There are 17,091,657 images and 360,232 ids. All images are aligned based on facial landmarks predicted by RetinaFace and resized to 112x112.
This was downloaded from https://github.com/deepinsight/insightface/tree/master/recognition/_datasets_. The original dataset format is MXNet RecordIO. It was converted to WebDataset in this copy here. There are 1,385 shards in total.… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/glint360k-wds-gz.csd_filesffhq-1024-wds
Flickr-Faces-HQ Dataset (FFHQ) - 1024x1024
This is a reupload of FFHQ-1024. Refer to the original dataset repo for more information https://github.com/NVlabs/ffhq-dataset
Specifically, this is the images1024x1024 set - faces are aligned and cropped to 1024x1024. Original PNG files were transcoded to WEBP losslessly to save space and packed to WebDataset format for ease of streaming. The original filenames are kept (with different file extension) so that you can match against… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/ffhq-1024-wds.laions_got_talent_rawgptsovits_dataset
bhyuan/gptsovits_dataset
GPT-SoVITS speech dataset, packed as WebDataset tar shards.
Layout
data/
train/
metadata.csv
audio/
train-000.tar
train-001.tar
...
validation/
metadata.csv
audio/
validation-000.tar
...
test/
metadata.csv
audio/
test-000.tar
...
Shard counts:
youshengshu_v5_test: 6536 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/bhyuan/gptsovits_dataset.voxceleb2-dev-wds
VoxCeleb2 - dev set
This is a copy of VoxCeleb2 dev set in WebDataset format. The audio data is the original AAC-encoded files without any transcoding. Refer to https://arxiv.org/abs/1806.05622 for more details about the dataset.
There are 1,092,009 samples covering 5,994 unique speakers. The dataset is split into 779 shards of ~100MB.
Usage
import torchaudio
import webdataset as wds
from datasets import load_dataset
def decode_audio(sample):
audio, fs =… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/voxceleb2-dev-wds.Hunyuan3D-FLUX-Gen
Orient Anything V2 Dataset
Project Page | Paper | GitHub
Orient Anything V2 is an enhanced foundation model for unified understanding of object 3D orientation and rotation from single or paired images. This dataset repository supports the model by providing assets for orientation estimation, 6DoF pose estimation, and object symmetry recognition.
Data Preparation
You can download the absolute orientation, relative rotation, and symm-orientation test datasets using the… See the full description on the dataset page: https://huggingface.co/datasets/Viglong/Hunyuan3D-FLUX-Gen.ms1mv3-wds
MS-Celeb-1M (v3)
This dataset is introduced in the Lightweight Face Recognition Challenge at ICCV 2019. Paper.
There are 5,179,510 images and 93,431 ids. All images are aligned based on facial landmarks predicted by RetinaFace and resized to 112x112.
This was downloaded from https://github.com/deepinsight/insightface/tree/master/recognition/_datasets_ (MS1M-RetinaFace). The original dataset format is MXNet RecordIO. It was converted to WebDataset in this copy here. There are 100… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/ms1mv3-wds.gs-videos-v3font-square-pretrain-20M
📚 Citation
If you use this dataset in your research, please cite these papers:
@article{pippi2023evaluating,
title={Evaluating Synthetic Pre-Training for Handwriting Processing Tasks},
author={Pippi, Vittorio and Cascianelli, Silvia and Baraldi, Lorenzo and Cucchiara, Rita},
journal={Pattern Recognition Letters},
year={2023},
publisher={Elsevier}
}
@InProceedings{pippi2025zeroshot,
author = {Pippi, Vittorio and Quattrini, Fabio and Cascianelli, Silvia and Tonioni… See the full description on the dataset page: https://huggingface.co/datasets/blowing-up-groundhogs/font-square-pretrain-20M.voice-data
Voice-Data: a curated multi-corpus voice dataset for voice–text contrastive training
voice-data is a single, globally-shuffled WebDataset that bundles several voice/speech corpora into one ready-to-train mixture for voice–text contrastive (CLAP-style) models such as VoiceCLAP. Each clip pairs 48 kHz mono FLAC audio with a natural-language text caption describing the voice — its emotion, prosody, timbre, speaking style, recording context, and speaker traits.
The distinguishing… See the full description on the dataset page: https://huggingface.co/datasets/gijs/voice-data.moby-gamesversion https://git-lfs.github.com/spec/v1
oid sha256:d8d7a46d41a1a37fe4f0a5f637bf55c649310185329127d8a2204632e480be17
size 24
Glint360k
Dataset Card for Glint360K
Citiation by InsightFace Repository
We clean, merge, and release the largest and cleanest face recognition dataset Glint360K, which contains 17091657 images of 360232 individuals. By employing the Patial FC training strategy, baseline models trained on Glint360K can easily achieve state-of-the-art performance. Detailed evaluation results on the large-scale test set (e.g. IFRT, IJB-C and Megaface) are as follows:
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yayoimizuha/Glint360k.dan-webp-newlaions_got_talent_enhanced_no_metadataGenRef-wds
GenRef-1M
We provide 1M high-quality triplets of the form (flawed image, high-quality image, reflection) collected across
multiple domains using our scalable pipeline from [1]. We used this dataset to train our reflection tuning model.
To know the details of the dataset creation pipeline, please refer to Section 3.2 of [1].
Project Page: https://diffusion-cot.github.io/reflection2perfection
Dataset loading
We provide the dataset in the webdataset format for fast… See the full description on the dataset page: https://huggingface.co/datasets/diffusion-cot/GenRef-wds.font-square-v2
Accessing the font-square-v2 Dataset on Hugging Face
The font-square-v2 dataset is hosted on Hugging Face at blowing-up-groundhogs/font-square-v2. It is stored in WebDataset format, with tar files organized as follows:
tars/train/: Contains {000..499}.tar shards for the main training split.
tars/fine_tune/: Contains {000..049}.tar shards for fine-tuning.
Each tar file contains multiple samples, where each sample includes:
An RGB image (.rgb.png)
A black-and-white image (.bw.png)… See the full description on the dataset page: https://huggingface.co/datasets/blowing-up-groundhogs/font-square-v2.mint-1t-html-images-gte6-sample
Size: 6769158 images sampled from Mint-1t-html
Criteria: Data entries with greater than or equal to 6 images (gte6)
laions_got_talent_embs_only
laions_got_talent Whisper Embeddings (Embeddings + Metadata Only)
This dataset contains Whisper embeddings (NPY) and metadata (JSON). The original audio files (MP3) are NOT included.
Embeddings computed with: mkrausio/EmoWhisper-AnS-Small-v0.1
Includes original audio: No
Includes metadata: Yes (JSON)
Includes embeddings: Yes (NPY)
Creation date: 2025-05-11
laions_got_talent_german_bicodecGalgame_Speech_SER_16kHz
Dataset Card for Galgame_Speech_SER_16kHz
[!IMPORTANT]The following rules (in the original repository) must be followed:
必须遵守GNU General Public License v3.0内的所有协议!附加:禁止商用,本数据集以及使用本数据集训练出来的任何模型都不得用于任何商业行为,如要用于商业用途,请找数据列表内的所有厂商授权(笑),因违反开源协议而出现的任何问题都与本人无关!
训练出来的模型必须开源,是否在README内引用本数据集由训练者自主决定,不做强制要求。
English:
You must comply with all the terms of the GNU General Public License v3.0!Additional note: Commercial use is prohibited. This dataset and any model trained using this dataset… See the full description on the dataset page: https://huggingface.co/datasets/litagin/Galgame_Speech_SER_16kHz.grit_2msdxl_images_easy_prompts-artists-seed1GlobalGeoTree
GlobalGeoTree Dataset
GlobalGeoTree is a comprehensive global dataset for tree species classification, comprising 6.3 million geolocated tree occurrences spanning 275 families, 2,734 genera, and 21,001 species across hierarchical taxonomic levels. Each sample is paired with Sentinel-2 image time series and 27 auxiliary environmental variables.
Dataset Structure
This repository contains three main components:
1. GlobalGeoTree-6M
Training dataset with around 6M… See the full description on the dataset page: https://huggingface.co/datasets/yann111/GlobalGeoTree.wikipedia_ru
