datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kasem-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Kasem Speech-Text Parallel Dataset
Dataset Description
This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kasem-speech-text-parallel.ga-speech-text-parallel-90k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Ga Speech-Text Parallel Dataset
Dataset Description
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ga-speech-text-parallel-90k.twi-trigrams-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-trigrams-speech-text-parallel.yoruba-speech-text-parallel
Yoruba Speech-Text Parallel Dataset
Dataset Description
This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in Nigeria and other West African countries. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Yoruba - yo
Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/yoruba-speech-text-parallel.vagla-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Vagla Speech-Text Parallel Dataset
Dataset Description
This dataset contains 48605 parallel speech-text pairs for Vagla, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/vagla-speech-text-parallel.makhuwa-trigrams-speech-text-parallel
Makhuwa Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 154253 parallel speech-text pairs for Makhuwa, a language spoken primarily in Mozambique. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Makhuwa - vmw
Task: Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/makhuwa-trigrams-speech-text-parallel.twi-words-speech-text-parallel-400k
Twi Words Speech-Text Parallel Dataset
Dataset Description
This dataset contains 413463 parallel speech-text pairs for Twi (Akan), a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Twi (Akan) - tw
Task: Speech Recognition, Text-to-Speech
Size: 413463 audio files >… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/twi-words-speech-text-parallel-400k.vai-speech-text-parallel
Vai Speech-Text Parallel Dataset
Dataset Description
This dataset contains 23286 parallel speech-text pairs for Vai, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Vai - vai
Task: Speech Recognition, Text-to-Speech
Size: 23286 audio files > 1KB (small/corrupted… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/vai-speech-text-parallel.deg-speech-text-parallel
Deg Speech-Text Parallel Dataset
Dataset Description
This dataset contains 125958 parallel speech-text pairs for Deg, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Deg - mzw
Task: Speech Recognition, Text-to-Speech
Size: 125958 audio files > 1KB (small/corrupted… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/deg-speech-text-parallel.swahili-words-speech-text-parallel
Swahili Words Speech-Text Parallel Dataset
Dataset Description
This dataset contains 411048 parallel speech-text pairs for Swahili, a widely spoken language in East Africa. The dataset consists of audio recordings paired with corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Swahili - sw
Task: Speech Recognition, Text-to-Speech
Size: 411048 audio files > 1KB… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-words-speech-text-parallel.synthetic-parallel-external
Synthetic Parallel EN↔LG — external
Voice-controlled synthetic parallel speech dataset for Luganda-English
speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline.
Generation
Component
Model
Translation
Sunbird/translate-nllb-3.3b-salt
TTS
Sunbird/orpheus-3b-tts-multilingual
English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003
Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003, waxal_lug_0004… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-external.paralingua_ru
Russian Paralinguistic Annotation Dataset
Датасет паралингвистической разметки спикеров из трёх русскоязычных корпусов:
biggest_ru_book,
DeepSpeech и Golos.
Что размечалось
Каждое аудио размечалось вручную по следующим характеристикам:
Поле
Описание
Пример значений
gender
Пол спикера
мужской, женский
age_group
Возрастная группа
молодой, взрослый, пожилой
voice_pitch
Высота голоса
низкий, средний, высокий
loudness
Громкость
тихий, нормальный… See the full description on the dataset page: https://huggingface.co/datasets/turnipseason/paralingua_ru.twi-trigrams-speech-text-parallel
Twi Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Twi - twi
Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/twi-trigrams-speech-text-parallel.chichewa-trigrams-speech-text-parallel
Chichewa Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 132549 parallel speech-text pairs for Chichewa, a language spoken primarily in Malawi. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Chichewa - ny
Task: Speech Recognition… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/chichewa-trigrams-speech-text-parallel.neuro-parakeet-food
neuro-whisper-v1
Dataset Description
This is a synthetic dataset for German medical speech recognition, specifically designed for fine-tuning ASR models on neuro-oncology and neurology terminology. The dataset provides a comprehensive coverage of German medical terminology in the neurology and neuro-oncology domains.
Data Generation
Voice Data: Synthetically generated using Resemble AI Chatterbox TTS
Text Data: Medical text generated with Qwen/Qwen3-30B-A3B… See the full description on the dataset page: https://huggingface.co/datasets/NeurologyAI/neuro-parakeet-food.synthetic-parallel-salt
Synthetic Parallel EN↔LG — salt
Voice-controlled synthetic parallel speech dataset for Luganda-English
speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline.
Generation
Component
Model
Translation
Sunbird/translate-nllb-3.3b-salt
TTS
Sunbird/orpheus-3b-tts-multilingual
English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003
Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003, waxal_lug_0004… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-salt.parakeet-tdt-blind-spots
Blind Spots of nvidia/parakeet-tdt-0.6b-v2
This dataset documents 14 systematically identified blind spots in NVIDIA's parakeet-tdt-0.6b-v2 automatic speech recognition model. The errors span 8 distinct categories and reveal a consistent pattern: the model struggles with inputs outside the distribution of its Western English-centric training data.
Model Under Test
Property
Value
Model
nvidia/parakeet-tdt-0.6b-v2
Parameters
600M
Architecture… See the full description on the dataset page: https://huggingface.co/datasets/TieIncred/parakeet-tdt-blind-spots.synthetic-parallel-sunbird
Synthetic Parallel EN↔LG — sunbird
Voice-controlled synthetic parallel speech dataset for Luganda-English
speech-to-speech translation, generated by the Hibiki-Zero fine-tuning pipeline.
Generation
Component
Model
Translation
Sunbird/translate-nllb-3.3b-salt
TTS
Sunbird/orpheus-3b-tts-multilingual
English speakers: salt_eng_0001, salt_eng_0002, salt_eng_0003
Luganda speakers: salt_lug_0001, waxal_lug_0001, waxal_lug_0002, waxal_lug_0003, waxal_lug_0004… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/synthetic-parallel-sunbird.lug-eng-synthetic-parallel-v1
Synthetic English-Luganda Parallel Speech (149,486 pairs)
Synthetic parallel audio for English-Luganda speech-to-speech translation
research, generated with Orpheus 3B TTS and stored as parquet shards with
embedded audio.
Schema
field
type
notes
id
string
pair id, e.g. pair_000123
audio_eng
Audio(22 050 Hz mono)
English clip
audio_lug
Audio(22 050 Hz mono)
Luganda clip
text_eng
string
English transcript
text_lug
string
Luganda transcript… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/lug-eng-synthetic-parallel-v1.
