CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01japanese-asr /whisper_transcriptions.reazon_speech_allaudio10M<n<100M16 likes47k downloads2y agoHugging Face02japanese-asr /whisper_transcriptions.reazonspeech.allaudio10M<n<100M4 likes7.1k downloads2y agoHugging Face03japanese-asr /whisper_transcriptions.reazonspeech.all.wer_10.0audio1M<n<10M3 likes3.5k downloads2y agoHugging Face04ARTPARK-IISc /Vaani-transcription-partgatedThis dataset is part of the Vaani dataset and consists of only transcribed speech data. It has a total duration of 2041.54 hours, covering 59 languages. This table represents the audio and transcription duration data for various languages. Language Angami Angika Ao Assamese Awadhi Bajjika Bearybashe Bengali Bhili Bhojpuri Bundeli Chakhesang Chakma Chhattisgarhi English Garhwali Garo Gondi Gujarati Halbi Haryanvi Hindi IduMishmi Kannada Kashmiri Karbi Khariboli Khortha Kokborok Konkani… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-transcription-part.audioautomatic-speech-recognition1M<n<10M20 likes1.6k downloads6mo agoHugging Face05jamescalam /youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp. Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.tabularquestion-answering100K<n<1M44 likes1.3k downloads4y agoHugging Face06japanese-asr /whisper_transcriptions.mlsaudio10M<n<100M1 likes927 downloads2y agoHugging Face07ghananlpcommunity /kumawood-speech-transcriptions Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/kumawood-speech-transcriptions.audioautomatic-speech-recognition100K<n<1M0 likes758 downloads22d agoHugging Face08japanese-asr /whisper_transcriptions.reazonspeech.large.wer_10.0audio1M<n<10M0 likes665 downloads3y agoHugging Face09ghanaopenai /kumawood-speech-transcriptions Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kumawood-speech-transcriptions.audioautomatic-speech-recognition100K<n<1M0 likes596 downloads22d agoHugging Face10japanese-asr /whisper_transcriptions.reazonspeech.mediumaudio100K<n<1M0 likes450 downloads3y agoHugging Face11openbank-uz /youtube_transcriptions Dataset Description A speech dataset of Uzbek language audio clips sourced from YouTube videos. Audio segments were extracted, separated by speaker using vocal isolation, and transcribed using Google's Gemini 2.0 Flash model. Speaker identities were clustered using ECAPA-TDNN embeddings. Use Cases Automatic Speech Recognition (ASR) for Uzbek Text-to-Speech (TTS) synthesis for Uzbek Fine-tuning speech models on Uzbek language data (e.g., Qwen3-TTS) Speaker-conditioned TTS… See the full description on the dataset page: https://huggingface.co/datasets/openbank-uz/youtube_transcriptions.audioautomatic-speech-recognition100K<n<1M2 likes443 downloads6mo agoHugging Face12octava /indonesian-voice-transcription-1.4.9a-raudio10K<n<100K0 likes406 downloads2y agoHugging Face13japanese-asr /whisper_transcriptions.reazonspeech.largeaudio1M<n<10M0 likes384 downloads3y agoHugging Face14octava /indonesian-voice-transcription-1.4.9a.2audio10K<n<100K0 likes378 downloads2y agoHugging Face15modulate /entity-transcription-benchmark Entity Transcription Benchmark Measures whether a speech recognition system transcribes named entities correctly — as distinct from word error rate. WER weights every token equally. The tokens that matter for redaction, lookup, routing and search are proper nouns, and they are a small fraction of any transcript. A system can improve WER while getting worse at exactly the words a downstream consumer needs, and nothing in the standard evaluation will show it. 2,151 clips, 6.0… See the full description on the dataset page: https://huggingface.co/datasets/modulate/entity-transcription-benchmark.audioautomatic-speech-recognition1K<n<10K4 likes336 downloads14d agoHugging Face16mugezhang /pair_tamil_malayalam_ipa_transcription_romanizedtext10M<n<100M0 likes277 downloads9mo agoHugging Face17octava /indonesian-voice-transcription-1.4.85aaudio10K<n<100K0 likes254 downloads2y agoHugging Face18octava /indonesian-voice-transcription-1.4.9a-cv-fl-slrjv-mdaudio10K<n<100K0 likes253 downloads2y agoHugging Face19octava /indonesian-voice-transcription-1.3.9caudio10K<n<100K0 likes246 downloads2y agoHugging Face20octava /indonesian-voice-transcription-1.4.9rcvaudio10K<n<100K0 likes246 downloads2y agoHugging Face21octava /indonesian-voice-transcription-1.4.9aaudio100K<n<1M0 likes241 downloads2y agoHugging Face22POTOMITAN /potomitan-gcf-transcriptiongated Kreyol Guadeloupe Transcription Dataset Ce jeu de données contient des segments audio courts (~5 secondes) en créole guadeloupéen (gcf), extraits d’émissions de radio et de télévision. Il vise à entraîner des modèles de reconnaissance automatique de la parole (ASR) pour une langue vivante mais peu disposant de peu de ressources écrites. Dataset Description Le créole guadeloupéen (Karukéya) est une langue créole à base lexicale française, parlée principalement en… See the full description on the dataset page: https://huggingface.co/datasets/POTOMITAN/potomitan-gcf-transcription.audioautomatic-speech-recognition10K<n<100K4 likes225 downloads10mo agoHugging Face23united-nations /transcription-corpus UN Transcription Corpus Two splits of UN meeting audio paired with official verbatim records. Splits sessions — Whole meeting sessions (SC + GA plenary) One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org. Column Description symbol UN document symbol, e.g. S/PV.9826 webtv_url URL on UN Web TV duration_ms Session duration in milliseconds num_speakers Number of speaker turns in the verbatim record audio_floor Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.audioautomatic-speech-recognitionn<1K0 likes222 downloads7mo agoHugging Face24octava /indonesian-voice-transcription-1.5.9a-cv-fl-slrjv-mdaudio100K<n<1M5 likes220 downloads2y agoHugging Face25jq /salt-asr-data-transcriptionstabular10K<n<100K0 likes208 downloads2y agoHugging Face26distil-whisper /whisper_transcriptions_greedytext10M<n<100M0 likes175 downloads3y agoHugging Face27octava /indonesian-voice-transcription-1.4.9raudio10K<n<100K0 likes175 downloads2y agoHugging Face28octava /indonesian-voice-transcription-1.0audio10K<n<100K3 likes155 downloads2y agoHugging Face29AdoCleanCode /correct_transcription_alignedaudio100K<n<1M0 likes151 downloads8mo agoHugging Face30tech4humans /Audio-Transcription-Models-Comparison-PT-BR Audio Transcription Models Comparison A dataset dedicated to comparing the performance of modern Speech-to-Text (STT) models, focusing exclusively on Brazilian Portuguese. About the Dataset This dataset was created to store and compare transcription results from different Artificial Intelligence models in challenging scenarios. Unlike generic benchmarks, this project focuses on the reality of usage in Brazil, covering: Regionalism: Local vocabulary, accents, and… See the full description on the dataset page: https://huggingface.co/datasets/tech4humans/Audio-Transcription-Models-Comparison-PT-BR.audioautomatic-speech-recognitionn<1K3 likes144 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.