datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
emolia-3k-speaker-clusters
Emolia 3K Speaker Clusters
A curated set of 3,000 diverse speaker clusters derived from the TTS-AGI/emolia-hq dataset, with up to 20 representative audio samples per cluster.
Overview
The original emolia-hq dataset contains hundreds of thousands of speech samples with 128-dimensional WavLM speaker timbre embeddings. These were first clustered into 10,000 centroids, then intelligently pruned to 3,000 using density-aware farthest-point sampling to ensure:
Outlier… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-3k-speaker-clusters.slurp_clustered_datasetjenny-tts-with-text-clustersvoice-gender-clustering
Dataset Details
VoxCelebs Dataset separated by gender (https://dagshub.com/DagsHub/audio-datasets/src/main/voice_gender_detection)
Dataset Description
Celebrities voice recordings separated by their gender.
Dataset Sources [optional]
VoxCeleb dataset (https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox2.html)
\Separation (https://dagshub.com/DagsHub/audio-datasets/src/main/voice_gender_detection)
slurp_clustered_split_dataset_fold1voxpopuli-accent-clusteringComplexly-cluster-thresh0_75-conf-thresh0_9Thai-Voice-Test-Clustering
Thanarit/Thai-Voice
Combined Thai audio dataset from multiple sources
Dataset Details
Total samples: 20
Total duration: 0.02 hours
Language: Thai (th)
Audio format: 16kHz mono WAV
Volume normalization: -20dB
Sources
Processed 1 datasets in streaming mode
Source Datasets
GigaSpeech2: Large-scale multilingual speech corpus
Usage
from datasets import load_dataset
# Load with streaming to avoid downloading everything
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Thanarit/Thai-Voice-Test-Clustering.sada_female_cluster2Thai-Voice-Test-Clustering-batch
Thanarit/Thai-Voice
Combined Thai audio dataset from multiple sources
Dataset Details
Total samples: 100
Total duration: 0.11 hours
Language: Thai (th)
Audio format: 16kHz mono WAV
Volume normalization: -20dB
Sources
Processed 1 datasets in streaming mode
Source Datasets
GigaSpeech2: Large-scale multilingual speech corpus
Usage
from datasets import load_dataset
# Load with streaming to avoid downloading everything
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Thanarit/Thai-Voice-Test-Clustering-batch.gol-dala-cluster
GOL voice clusters — audited repair
This audit covers midralab/gol-dala-cluster revision
89e1c5982086207f0a8de22cdb880e5ea52789f6. The original repository has no dataset
card and stores its files under Windows-style backslash paths. Its 89-byte
cluster_statistics.json ends inside the cluster_statistics object and is invalid
JSON.
Verified source layout
voice_clusters.csv: 7,362,684 strictly parsed rows, 596 folders, 19,245 speakers,
150 cluster IDs (0–149), and… See the full description on the dataset page: https://huggingface.co/datasets/midralab/gol-dala-cluster.
