datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
whisper_transcriptions.reazon_speech_allwhisper_transcriptions.reazonspeech.allwhisper_transcriptions.reazonspeech.all.wer_10.0Vaani-transcription-partThis dataset is part of the Vaani dataset and consists of only transcribed speech data. It has a total duration of 2041.54 hours, covering 59 languages.
This table represents the audio and transcription duration data for various languages.
Language
Angami
Angika
Ao
Assamese
Awadhi
Bajjika
Bearybashe
Bengali
Bhili
Bhojpuri
Bundeli
Chakhesang
Chakma
Chhattisgarhi
English
Garhwali
Garo
Gondi
Gujarati
Halbi
Haryanvi
Hindi
IduMishmi
Kannada
Kashmiri
Karbi
Khariboli
Khortha
Kokborok
Konkani… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-transcription-part.youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp.
Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.whisper_transcriptions.mlskumawood-speech-transcriptions
Kumawood Speech Transcriptions
Speech segments from Ghanaian films, each paired with the film's human-authored
English subtitle and a machine Twi transcript.
Total number of hours 249.8 hours
Fields
field
meaning
audio
16 kHz mono FLAC segment
text
English subtitle displayed during the segment (human-authored, recovered by OCR)
twi_text
Twi transcript from Google STT (ak) — machine output
twi_words_per_sec
transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/kumawood-speech-transcriptions.whisper_transcriptions.reazonspeech.large.wer_10.0kumawood-speech-transcriptions
Kumawood Speech Transcriptions
Speech segments from Ghanaian films, each paired with the film's human-authored
English subtitle and a machine Twi transcript.
Total number of hours 249.8 hours
Fields
field
meaning
audio
16 kHz mono FLAC segment
text
English subtitle displayed during the segment (human-authored, recovered by OCR)
twi_text
Twi transcript from Google STT (ak) — machine output
twi_words_per_sec
transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kumawood-speech-transcriptions.whisper_transcriptions.reazonspeech.mediumyoutube_transcriptions
Dataset Description
A speech dataset of Uzbek language audio clips sourced from YouTube videos. Audio segments were extracted, separated by speaker using vocal isolation, and transcribed using Google's Gemini 2.0 Flash model. Speaker identities were clustered using ECAPA-TDNN embeddings.
Use Cases
Automatic Speech Recognition (ASR) for Uzbek
Text-to-Speech (TTS) synthesis for Uzbek
Fine-tuning speech models on Uzbek language data (e.g., Qwen3-TTS)
Speaker-conditioned TTS… See the full description on the dataset page: https://huggingface.co/datasets/openbank-uz/youtube_transcriptions.indonesian-voice-transcription-1.4.9a-rwhisper_transcriptions.reazonspeech.largeindonesian-voice-transcription-1.4.9a.2entity-transcription-benchmark
Entity Transcription Benchmark
Measures whether a speech recognition system transcribes named entities
correctly — as distinct from word error rate.
WER weights every token equally. The tokens that matter for redaction, lookup,
routing and search are proper nouns, and they are a small fraction of any
transcript. A system can improve WER while getting worse at exactly the words a
downstream consumer needs, and nothing in the standard evaluation will show it.
2,151 clips, 6.0… See the full description on the dataset page: https://huggingface.co/datasets/modulate/entity-transcription-benchmark.pair_tamil_malayalam_ipa_transcription_romanizedindonesian-voice-transcription-1.4.85aindonesian-voice-transcription-1.4.9a-cv-fl-slrjv-mdindonesian-voice-transcription-1.3.9cindonesian-voice-transcription-1.4.9rcvindonesian-voice-transcription-1.4.9apotomitan-gcf-transcription
Kreyol Guadeloupe Transcription Dataset
Ce jeu de données contient des segments audio courts (~5 secondes) en créole guadeloupéen (gcf), extraits d’émissions de radio et de télévision.
Il vise à entraîner des modèles de reconnaissance automatique de la parole (ASR) pour une langue vivante mais peu disposant de peu de ressources écrites.
Dataset Description
Le créole guadeloupéen (Karukéya) est une langue créole à base lexicale française, parlée principalement en… See the full description on the dataset page: https://huggingface.co/datasets/POTOMITAN/potomitan-gcf-transcription.transcription-corpus
UN Transcription Corpus
Two splits of UN meeting audio paired with official verbatim records.
Splits
sessions — Whole meeting sessions (SC + GA plenary)
One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org.
Column
Description
symbol
UN document symbol, e.g. S/PV.9826
webtv_url
URL on UN Web TV
duration_ms
Session duration in milliseconds
num_speakers
Number of speaker turns in the verbatim record
audio_floor
Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.indonesian-voice-transcription-1.5.9a-cv-fl-slrjv-mdsalt-asr-data-transcriptionswhisper_transcriptions_greedyindonesian-voice-transcription-1.4.9rindonesian-voice-transcription-1.0correct_transcription_alignedAudio-Transcription-Models-Comparison-PT-BR
Audio Transcription Models Comparison
A dataset dedicated to comparing the performance of modern Speech-to-Text (STT) models, focusing exclusively on Brazilian Portuguese.
About the Dataset
This dataset was created to store and compare transcription results from different Artificial Intelligence models in challenging scenarios. Unlike generic benchmarks, this project focuses on the reality of usage in Brazil, covering:
Regionalism: Local vocabulary, accents, and… See the full description on the dataset page: https://huggingface.co/datasets/tech4humans/Audio-Transcription-Models-Comparison-PT-BR.
