datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
test-resumeThai-Voice-Test-Resume
Thanarit/Thai-Voice
Combined Thai audio dataset from multiple sources
Dataset Details
Total samples: 200
Total duration: 0.22 hours
Language: Thai (th)
Audio format: 16kHz mono WAV
Volume normalization: -20dB
Sources
Processed 1 datasets in streaming mode
Source Datasets
GigaSpeech2: Large-scale multilingual speech corpus
Usage
from datasets import load_dataset
# Load with streaming to avoid downloading everything
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Thanarit/Thai-Voice-Test-Resume.
