datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gigaspeech2
Dataset Card for GigaSpeech 2
Dataset Description
GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese.
Repository: https://github.com/SpeechColab/GigaSpeech2
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2.data_audio_gigaspeech2_Educationdata_audio_gigaspeech2_Entertainmentdata_audio_gigaspeech2_Education_copy1gigaspeech2-th-joined
gigaspeech2-th-joined
Thai speech-transcript pairs derived from GigaSpeech 2,
re-segmented into clips of a length that is convenient for training speech-language models.
Built to train a Thai speech adapter for the Ultravox
architecture, where very short fragments make poor training examples but long clips do not fit
the context budget.
Statistics
Examples
40,000
Total audio
~42 hours
Sample rate
16 kHz, mono
Clip duration
2.0–12.0 s (mean 3.8… See the full description on the dataset page: https://huggingface.co/datasets/Funk888/gigaspeech2-th-joined.thai_gigaspeech2Thai part of gigaspeech2: https://huggingface.co/datasets/speechcolab/gigaspeech2
gigaspeech2-test@article{yang2024gigaspeech,
title={GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement},
author={Yang, Yifan and Song, Zheshu and Zhuo, Jianheng and Cui, Mingyu and Li, Jinpeng and Yang, Bo and Du, Yexing and Ma, Ziyang and Liu, Xunying and Wang, Ziyuan and others},
journal={arXiv preprint arXiv:2406.11546},
year={2024}
}
@article{wang2024audiobench,
title={AudioBench: A Universal… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/gigaspeech2-test.gigaspeech2-typhoon
Gigaspeech2 Typhoon
Project page | Paper | GitHub
Gigaspeech2 Typhoon is a metadata-only reference dataset for Thai speech recognition benchmarking, specifically designed as an Accuracy Track for evaluating ASR models. The dataset contains 1,000 test samples with audio IDs and human transcriptions derived from the Gigaspeech2 corpus. Each audio_id directly links to the original Gigaspeech2 dataset, allowing users to download the corresponding audio.
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/gigaspeech2-typhoon.gigaspeech2_metadatadata_audio_gigaspeech2gigaspeech2_vie
Vietnamse subset of the Gigaspeech2 dataset
extracted from: https://huggingface.co/datasets/speechcolab/gigaspeech2
gigaspeech2-test
Dataset Card for GigaSpeech 2 TEST
Dataset Description
GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese.
Repository: https://github.com/SpeechColab/GigaSpeech2
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2-test.data_audio_gigaspeech2_Education_copygigaspeech2-vi-missing-transcripts
GigaSpeech2 Vietnamese WAVs Missing Transcripts
This dataset contains the Vietnamese GigaSpeech2 training WAV files whose IDs
are absent from train_refined.tsv.
Contents
201,295 WAV files without a matching transcript
41 uncompressed TAR shards in shards/
36 GB of audio (approximately)
processing_manifest.jsonl: per-source-archive counts
summary.json: aggregate counts
SHA256SUMS: checksums for all TAR shards
Each shard preserves the original relative layout:… See the full description on the dataset page: https://huggingface.co/datasets/quangdung/gigaspeech2-vi-missing-transcripts.gigaspeech2_test_with_noisegigaspeech2-th-joined-v2gigaspeech2_DEMANDgigaspeech2_denoisedata_audio_gigaspeech2preprocessed_gigaspeech2gigaspeech2_borroweddata_audio_gigaspeech2_Fitnessdata_audio_gigaspeech2_Testinggigaspeech2-id-parquet
