CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01speechcolab /gigaspeech2gated Dataset Card for GigaSpeech 2 Dataset Description GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese. Repository: https://github.com/SpeechColab/GigaSpeech2 Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2.audioautomatic-speech-recognition10M<n<100M71 likes5.8k downloads6mo agoHugging Face02tranvy /data_audio_gigaspeech2_Educationaudio1K<n<10K0 likes407 downloads1y agoHugging Face03tranvy /data_audio_gigaspeech2_Entertainmentaudio10K<n<100K0 likes324 downloads1y agoHugging Face04tranvy /data_audio_gigaspeech2_Education_copy1audio10K<n<100K0 likes306 downloads1y agoHugging Face05Funk888 /gigaspeech2-th-joined gigaspeech2-th-joined Thai speech-transcript pairs derived from GigaSpeech 2, re-segmented into clips of a length that is convenient for training speech-language models. Built to train a Thai speech adapter for the Ultravox architecture, where very short fragments make poor training examples but long clips do not fit the context budget. Statistics Examples 40,000 Total audio ~42 hours Sample rate 16 kHz, mono Clip duration 2.0–12.0 s (mean 3.8… See the full description on the dataset page: https://huggingface.co/datasets/Funk888/gigaspeech2-th-joined.textautomatic-speech-recognition10K<n<100K1 likes119 downloads17d agoHugging Face06huangjianbyte2024 /thai_gigaspeech2Thai part of gigaspeech2: https://huggingface.co/datasets/speechcolab/gigaspeech2 audio1K<n<10K3 likes106 downloads2y agoHugging Face07AudioLLMs /gigaspeech2-test@article{yang2024gigaspeech, title={GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement}, author={Yang, Yifan and Song, Zheshu and Zhuo, Jianheng and Cui, Mingyu and Li, Jinpeng and Yang, Bo and Du, Yexing and Ma, Ziyang and Liu, Xunying and Wang, Ziyuan and others}, journal={arXiv preprint arXiv:2406.11546}, year={2024} } @article{wang2024audiobench, title={AudioBench: A Universal… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/gigaspeech2-test.audio10K<n<100K1 likes103 downloads2y agoHugging Face08typhoon-ai /gigaspeech2-typhoon Gigaspeech2 Typhoon Project page | Paper | GitHub Gigaspeech2 Typhoon is a metadata-only reference dataset for Thai speech recognition benchmarking, specifically designed as an Accuracy Track for evaluating ASR models. The dataset contains 1,000 test samples with audio IDs and human transcriptions derived from the Gigaspeech2 corpus. Each audio_id directly links to the original Gigaspeech2 dataset, allowing users to download the corresponding audio. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/gigaspeech2-typhoon.textautomatic-speech-recognition1K<n<10K1 likes60 downloads4mo agoHugging Face09zzasdf /gigaspeech2_metadata0 likes37 downloads2y agoHugging Face10tranvy /data_audio_gigaspeech2audion<1K0 likes34 downloads2y agoHugging Face11doof-ferb /gigaspeech2_vie Vietnamse subset of the Gigaspeech2 dataset extracted from: https://huggingface.co/datasets/speechcolab/gigaspeech2 audioautomatic-speech-recognition10K<n<100K1 likes32 downloads1y agoHugging Face12speechcolab /gigaspeech2-test Dataset Card for GigaSpeech 2 TEST Dataset Description GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese. Repository: https://github.com/SpeechColab/GigaSpeech2 Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2-test.automatic-speech-recognition1M<n<10M0 likes26 downloads6mo agoHugging Face13tranvy /data_audio_gigaspeech2_Education_copyaudio1K<n<10K0 likes18 downloads1y agoHugging Face14quangdung /gigaspeech2-vi-missing-transcripts GigaSpeech2 Vietnamese WAVs Missing Transcripts This dataset contains the Vietnamese GigaSpeech2 training WAV files whose IDs are absent from train_refined.tsv. Contents 201,295 WAV files without a matching transcript 41 uncompressed TAR shards in shards/ 36 GB of audio (approximately) processing_manifest.jsonl: per-source-archive counts summary.json: aggregate counts SHA256SUMS: checksums for all TAR shards Each shard preserves the original relative layout:… See the full description on the dataset page: https://huggingface.co/datasets/quangdung/gigaspeech2-vi-missing-transcripts.audioautomatic-speech-recognition0 likes16 downloads2mo agoHugging Face15Duyynh /gigaspeech2_test_with_noise0 likes15 downloads1y agoHugging Face16Funk888 /gigaspeech2-th-joined-v2text100K<n<1M0 likes12 downloads4mo agoHugging Face17Duyynh /gigaspeech2_DEMANDtext10K<n<100K0 likes9 downloads1y agoHugging Face18Duyynh /gigaspeech2_denoiseaudio10K<n<100K0 likes7 downloads1y agoHugging Face19tranpoline /data_audio_gigaspeech2audion<1K0 likes6 downloads1y agoHugging Face20nhanv /preprocessed_gigaspeech20 likes6 downloads1y agoHugging Face21tuandung2812 /gigaspeech2_borrowedaudio1K<n<10K0 likes4 downloads9mo agoHugging Face22tranvy /data_audio_gigaspeech2_Fitnessaudio0 likes3 downloads1y agoHugging Face23tranvy /data_audio_gigaspeech2_Testing0 likes2 downloads1y agoHugging Face24ai4deafblind /gigaspeech2-id-parquet0 likes2 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.