datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voice-gender-clustering
Dataset Details
VoxCelebs Dataset separated by gender (https://dagshub.com/DagsHub/audio-datasets/src/main/voice_gender_detection)
Dataset Description
Celebrities voice recordings separated by their gender.
Dataset Sources [optional]
VoxCeleb dataset (https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox2.html)
\Separation (https://dagshub.com/DagsHub/audio-datasets/src/main/voice_gender_detection)
voxpopuli-accent-clusteringThai-Voice-Test-Clustering
Thanarit/Thai-Voice
Combined Thai audio dataset from multiple sources
Dataset Details
Total samples: 20
Total duration: 0.02 hours
Language: Thai (th)
Audio format: 16kHz mono WAV
Volume normalization: -20dB
Sources
Processed 1 datasets in streaming mode
Source Datasets
GigaSpeech2: Large-scale multilingual speech corpus
Usage
from datasets import load_dataset
# Load with streaming to avoid downloading everything
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Thanarit/Thai-Voice-Test-Clustering.Thai-Voice-Test-Clustering-batch
Thanarit/Thai-Voice
Combined Thai audio dataset from multiple sources
Dataset Details
Total samples: 100
Total duration: 0.11 hours
Language: Thai (th)
Audio format: 16kHz mono WAV
Volume normalization: -20dB
Sources
Processed 1 datasets in streaming mode
Source Datasets
GigaSpeech2: Large-scale multilingual speech corpus
Usage
from datasets import load_dataset
# Load with streaming to avoid downloading everything
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Thanarit/Thai-Voice-Test-Clustering-batch.
