gigaspeech2
gigaspeech2
Dataset Card for GigaSpeech 2
Dataset Description
GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese.
Repository: https://github.com/SpeechColab/GigaSpeech2
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2.data_audio_gigaspeech2_Educationdata_audio_gigaspeech2_Entertainmentdata_audio_gigaspeech2_Education_copy1gigaspeech2-th-joined
gigaspeech2-th-joined
Thai speech-transcript pairs derived from GigaSpeech 2,
re-segmented into clips of a length that is convenient for training speech-language models.
Built to train a Thai speech adapter for the Ultravox
architecture, where very short fragments make poor training examples but long clips do not fit
the context budget.
Statistics
Examples
40,000
Total audio
~42 hours
Sample rate
16 kHz, mono
Clip duration
2.0–12.0 s (mean 3.8… See the full description on the dataset page: https://huggingface.co/datasets/Funk888/gigaspeech2-th-joined.thai_gigaspeech2Thai part of gigaspeech2: https://huggingface.co/datasets/speechcolab/gigaspeech2
