datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kasem-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Kasem Speech-Text Parallel Dataset
Dataset Description
This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kasem-speech-text-parallel.ga-speech-text-parallel-90k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Ga Speech-Text Parallel Dataset
Dataset Description
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ga-speech-text-parallel-90k.twi-trigrams-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-trigrams-speech-text-parallel.yoruba-speech-text-parallel
Yoruba Speech-Text Parallel Dataset
Dataset Description
This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in Nigeria and other West African countries. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Yoruba - yo
Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/yoruba-speech-text-parallel.vagla-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Vagla Speech-Text Parallel Dataset
Dataset Description
This dataset contains 48605 parallel speech-text pairs for Vagla, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/vagla-speech-text-parallel.makhuwa-trigrams-speech-text-parallel
Makhuwa Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 154253 parallel speech-text pairs for Makhuwa, a language spoken primarily in Mozambique. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Makhuwa - vmw
Task: Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/makhuwa-trigrams-speech-text-parallel.twi-words-speech-text-parallel-400k
Twi Words Speech-Text Parallel Dataset
Dataset Description
This dataset contains 413463 parallel speech-text pairs for Twi (Akan), a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Twi (Akan) - tw
Task: Speech Recognition, Text-to-Speech
Size: 413463 audio files >… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/twi-words-speech-text-parallel-400k.vai-speech-text-parallel
Vai Speech-Text Parallel Dataset
Dataset Description
This dataset contains 23286 parallel speech-text pairs for Vai, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Vai - vai
Task: Speech Recognition, Text-to-Speech
Size: 23286 audio files > 1KB (small/corrupted… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/vai-speech-text-parallel.deg-speech-text-parallel
Deg Speech-Text Parallel Dataset
Dataset Description
This dataset contains 125958 parallel speech-text pairs for Deg, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Deg - mzw
Task: Speech Recognition, Text-to-Speech
Size: 125958 audio files > 1KB (small/corrupted… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/deg-speech-text-parallel.swahili-words-speech-text-parallel
Swahili Words Speech-Text Parallel Dataset
Dataset Description
This dataset contains 411048 parallel speech-text pairs for Swahili, a widely spoken language in East Africa. The dataset consists of audio recordings paired with corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Swahili - sw
Task: Speech Recognition, Text-to-Speech
Size: 411048 audio files > 1KB… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-words-speech-text-parallel.synthetic-parallel-external
Synthetic Parallel EN↔LG — external
Voice-controlled synthetic parallel speech dataset for Luganda-English
speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline.
Generation
Component
Model
Translation
Sunbird/translate-nllb-3.3b-salt
TTS
Sunbird/orpheus-3b-tts-multilingual
English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003
Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003, waxal_lug_0004… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-external.paralingua_ru
Russian Paralinguistic Annotation Dataset
Датасет паралингвистической разметки спикеров из трёх русскоязычных корпусов:
biggest_ru_book,
DeepSpeech и Golos.
Что размечалось
Каждое аудио размечалось вручную по следующим характеристикам:
Поле
Описание
Пример значений
gender
Пол спикера
мужской, женский
age_group
Возрастная группа
молодой, взрослый, пожилой
voice_pitch
Высота голоса
низкий, средний, высокий
loudness
Громкость
тихий, нормальный… See the full description on the dataset page: https://huggingface.co/datasets/turnipseason/paralingua_ru.twi-trigrams-speech-text-parallel
Twi Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Twi - twi
Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/twi-trigrams-speech-text-parallel.chichewa-trigrams-speech-text-parallel
Chichewa Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 132549 parallel speech-text pairs for Chichewa, a language spoken primarily in Malawi. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Chichewa - ny
Task: Speech Recognition… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/chichewa-trigrams-speech-text-parallel.fleurs-tr-hi-parallel-speech
FLEURS TR↔HI Parallel Speech
Turkish⇄Hindi parallel speech built from FLEURS — the real human speech
counterpart to this project's synthetic TTS corpora.
audio/
~8,935 clips
fleurs/
2,440 source FLEURS files
manifests/
selection + QC manifests (incl. accepted.jsonl)
Mimi-encoded downstream as
fleurs-tr-hi-mimi-encoded,
which is what the v0.3 evaluation actually consumed.
⚠️ Acoustic shift, not held-out text
An overlap audit of the derived… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/fleurs-tr-hi-parallel-speech.parakeet-hindi-asr
Parakeet Hindi-English Bilingual ASR
Fine-tuning NVIDIA Parakeet TDT 0.6B for bilingual Hindi-English automatic speech recognition.
Quick Start
# Download
pip install huggingface_hub
huggingface-cli download ketav/parakeet-hindi-asr --repo-type dataset --local-dir ./parakeet-hindi-asr
# Install dependencies
pip install nemo_toolkit[asr] bitsandbytes sentencepiece
# Train (after updating paths in config)
cd parakeet-hindi-asr/scripts
python ft_0.6B_hi_v3.py… See the full description on the dataset page: https://huggingface.co/datasets/ketav/parakeet-hindi-asr.neuro-parakeet-food
neuro-whisper-v1
Dataset Description
This is a synthetic dataset for German medical speech recognition, specifically designed for fine-tuning ASR models on neuro-oncology and neurology terminology. The dataset provides a comprehensive coverage of German medical terminology in the neurology and neuro-oncology domains.
Data Generation
Voice Data: Synthetically generated using Resemble AI Chatterbox TTS
Text Data: Medical text generated with Qwen/Qwen3-30B-A3B… See the full description on the dataset page: https://huggingface.co/datasets/NeurologyAI/neuro-parakeet-food.mile_datasetIISc-MILE Tamil ASR Corpus contains transcribed speech corpus for training ASR systems for Tamil language. It contains ~150 hours of read speech data collected from 531 speakers in a noise-free recording environment with high quality USB microphones.Indic_Hindi-English_Parallel_Speech
Dataset Access Information
This dataset is provided for research and academic purposes. Access to the dataset is gated, and users must request permission before downloading.
Dataset Summary
This repository contains the Hindi–English Speech-to-Speech Translation (S2ST) dataset introduced in the paper:
Benchmarking Hindi-to-English Direct Speech-to-Speech Translation with Synthetic Data
The dataset is designed to support research on direct speech-to-speech translation… See the full description on the dataset page: https://huggingface.co/datasets/mahendraphd/Indic_Hindi-English_Parallel_Speech.tamil_asr_corpusThe corpus contains roughly 1000 hours of audio and trasncripts in Tamil language. The transcripts have beedn de-duplicated using exact match deduplication.bengali_asr_corpusThe corpus contains roughly 500 hours of audio and transcripts in Bangla language.
The transcripts have beed de-duplicated using exact match deduplication and audio has be converted to 16000 samplessynthetic-parallel-salt
Synthetic Parallel EN↔LG — salt
Voice-controlled synthetic parallel speech dataset for Luganda-English
speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline.
Generation
Component
Model
Translation
Sunbird/translate-nllb-3.3b-salt
TTS
Sunbird/orpheus-3b-tts-multilingual
English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003
Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003, waxal_lug_0004… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-salt.kannada_asr_corpusThe corpus contains roughly 360 hours of audio and transcripts in Kannada language. The transcripts have beed de-duplicated using exact match deduplication.malayalam_asr_corpusThe corpus contains roughly 10 hours of audio and trasncripts in Malayalam language. The transcripts have beedn de-duplicated using exact match deduplication.bhasaflow-khasi-english-parallel-sample-v1
BhasaFlow Khasi-English Parallel Sample v1
A professionally curated, gold-standard parallel speech and text corpus for the Khasi language.
Published by Medharvix Systems Private Limited
Part of the BhasaFlow Low-Resource Language Technology Initiative
Overview
This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-sample-v1.parakeet-tdt-blind-spots
Blind Spots of nvidia/parakeet-tdt-0.6b-v2
This dataset documents 14 systematically identified blind spots in NVIDIA's parakeet-tdt-0.6b-v2 automatic speech recognition model. The errors span 8 distinct categories and reveal a consistent pattern: the model struggles with inputs outside the distribution of its Western English-centric training data.
Model Under Test
Property
Value
Model
nvidia/parakeet-tdt-0.6b-v2
Parameters
600M
Architecture… See the full description on the dataset page: https://huggingface.co/datasets/TieIncred/parakeet-tdt-blind-spots.bhasaflow-khasi-english-parallel-sample-v1
BhasaFlow Khasi-English Parallel Sample v1
A professionally curated, gold-standard parallel speech and text corpus for the Khasi language.
Published by Medharvix Systems Private Limited
Part of the BhasaFlow Low-Resource Language Technology Initiative
Overview
This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/1infinity0/bhasaflow-khasi-english-parallel-sample-v1.telugu_asr_corpusThe corpus contains roughly 360 hours of audio and transcripts in Telugu language. The transcripts have beed de-duplicated using exact match deduplication.cantonese-chinese-parallel-audio
Dataset Summary
This dataset aims to provide high-quality Cantonese (Yue) – Chinese (Zh) bilingual speech recordings with accurate transcriptions, designed for training and evaluating Automatic Speech Recognition (ASR) systems in multilingual and dialect-rich scenarios.
This dataset is currently under preparation and will be released soon.
Language
Cantonese (yue)
Chinese (zh)
Citation
Please check back later for the final publication details.
ucla_datasetTHE UCLA Tamil Labelled Total Duration contains 1160.24 hours of labelled ASR audio collected from various sources
