datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
X-Voice-TestsetX-Voice Multilingual Test Set
High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages
This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model.
Dataset Summary
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.UCF-Crime_TEST_SEThf download backseollgi/UCF-Crime_TEST_SET --repo-type dataset --local-dir .
tar -xzvf UCF-Crime_TEST_SET.tar.gz
Garments2Look-Test-Set-Results
Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories
Project Page | Paper | Code
Garments2Look is a large-scale multimodal dataset for outfit-level Virtual Try-On (VTON), comprising 80,000 many-garments-to-one-look pairs across 40 major categories and over 300 fine-grained subcategories. Each pair includes an outfit with 3-12 reference garment images (averaging 4.48), a model image wearing the outfit, and detailed item… See the full description on the dataset page: https://huggingface.co/datasets/ArtmeScienceLab/Garments2Look-Test-Set-Results.Garments2Look-Test-Set-Results
Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories
Project Page | Paper | Code
Garments2Look is a large-scale multimodal dataset for outfit-level Virtual Try-On (VTON), comprising 80,000 many-garments-to-one-look pairs across 40 major categories and over 300 fine-grained subcategories. Each pair includes an outfit with 3-12 reference garment images (averaging 4.48), a model image wearing the outfit, and detailed… See the full description on the dataset page: https://huggingface.co/datasets/baajarmah/Garments2Look-Test-Set-Results.test-setasc_testset
Yougen/asc_testset
Audio Scene Classification (ASC) speech dataset, packed as WebDataset tar shards.
Layout
data/
train/
metadata.csv
audio/
train-000.tar
train-001.tar
...
validation/
metadata.csv
audio/
validation-000.tar
...
test/
metadata.csv
audio/
test-000.tar
...
Shard counts:
test_a1: 8 tar shard(s)
test_a2: 16 tar shard(s)
test_a3: 13 tar shard(s)
test_a4: 15 tar shard(s)
test_a5: 8 tar… See the full description on the dataset page: https://huggingface.co/datasets/Yougen/asc_testset.said_testset
Yougen/said_testset
Speaker Recognition (SRE) speech dataset, packed as WebDataset tar shards.
Layout
data/
train/
metadata.csv
audio/
train-000.tar
train-001.tar
...
validation/
metadata.csv
audio/
validation-000.tar
...
test/
metadata.csv
audio/
test-000.tar
...
Shard counts:
test: 20 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw audio bytes… See the full description on the dataset page: https://huggingface.co/datasets/Yougen/said_testset.
