CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01speechcolab /gigaspeech2gated Dataset Card for GigaSpeech 2 Dataset Description GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese. Repository: https://github.com/SpeechColab/GigaSpeech2 Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2.audioautomatic-speech-recognition10M<n<100M71 likes5.8k downloads6mo agoHugging Face02tranvy /data_audio_gigaspeech2_Educationaudio1K<n<10K0 likes407 downloads1y agoHugging Face03tranvy /data_audio_gigaspeech2_Entertainmentaudio10K<n<100K0 likes324 downloads1y agoHugging Face04tranvy /data_audio_gigaspeech2_Education_copy1audio10K<n<100K0 likes306 downloads1y agoHugging Face05Funk888 /gigaspeech2-th-joined gigaspeech2-th-joined Thai speech-transcript pairs derived from GigaSpeech 2, re-segmented into clips of a length that is convenient for training speech-language models. Built to train a Thai speech adapter for the Ultravox architecture, where very short fragments make poor training examples but long clips do not fit the context budget. Statistics Examples 40,000 Total audio ~42 hours Sample rate 16 kHz, mono Clip duration 2.0–12.0 s (mean 3.8… See the full description on the dataset page: https://huggingface.co/datasets/Funk888/gigaspeech2-th-joined.textautomatic-speech-recognition10K<n<100K1 likes119 downloads18d agoHugging Face06huangjianbyte2024 /thai_gigaspeech2Thai part of gigaspeech2: https://huggingface.co/datasets/speechcolab/gigaspeech2 audio1K<n<10K3 likes106 downloads2y agoHugging Face07AudioLLMs /gigaspeech2-test@article{yang2024gigaspeech, title={GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement}, author={Yang, Yifan and Song, Zheshu and Zhuo, Jianheng and Cui, Mingyu and Li, Jinpeng and Yang, Bo and Du, Yexing and Ma, Ziyang and Liu, Xunying and Wang, Ziyuan and others}, journal={arXiv preprint arXiv:2406.11546}, year={2024} } @article{wang2024audiobench, title={AudioBench: A Universal… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/gigaspeech2-test.audio10K<n<100K1 likes103 downloads2y agoHugging Face08typhoon-ai /gigaspeech2-typhoon Gigaspeech2 Typhoon Project page | Paper | GitHub Gigaspeech2 Typhoon is a metadata-only reference dataset for Thai speech recognition benchmarking, specifically designed as an Accuracy Track for evaluating ASR models. The dataset contains 1,000 test samples with audio IDs and human transcriptions derived from the Gigaspeech2 corpus. Each audio_id directly links to the original Gigaspeech2 dataset, allowing users to download the corresponding audio. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/gigaspeech2-typhoon.textautomatic-speech-recognition1K<n<10K1 likes60 downloads4mo agoHugging Face09doof-ferb /gigaspeech2_vie Vietnamse subset of the Gigaspeech2 dataset extracted from: https://huggingface.co/datasets/speechcolab/gigaspeech2 audioautomatic-speech-recognition10K<n<100K1 likes32 downloads1y agoHugging Face10tranvy /data_audio_gigaspeech2_Education_copyaudio1K<n<10K0 likes18 downloads1y agoHugging Face11Funk888 /gigaspeech2-th-joined-v2text100K<n<1M0 likes12 downloads4mo agoHugging Face12Duyynh /gigaspeech2_DEMANDtext10K<n<100K0 likes9 downloads1y agoHugging Face13tranpoline /data_audio_gigaspeech2audion<1K0 likes6 downloads1y agoHugging Face14tuandung2812 /gigaspeech2_borrowedaudio1K<n<10K0 likes4 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.