datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twi-words-speech-text-parallel-400k
Twi Words Speech-Text Parallel Dataset
Dataset Description
This dataset contains 413463 parallel speech-text pairs for Twi (Akan), a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Twi (Akan) - tw
Task: Speech Recognition, Text-to-Speech
Size: 413463 audio files >… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/twi-words-speech-text-parallel-400k.swahili-words-speech-text-parallel
Swahili Words Speech-Text Parallel Dataset
Dataset Description
This dataset contains 411048 parallel speech-text pairs for Swahili, a widely spoken language in East Africa. The dataset consists of audio recordings paired with corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Swahili - sw
Task: Speech Recognition, Text-to-Speech
Size: 411048 audio files > 1KB… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-words-speech-text-parallel.persian-words
Persian Words
This is a dataset of approximately 5K words, read aloud by a variety of native speakers. The dataset has been directly redistributed from this URL.
It can be used as a valuable resource for evaluating/training ASR engines or speech synthesis engines.
P.S.: I'm not the original creator of this dataset, for crediting or ownership change you can contact
target-words-geminitts
Target-Word (TW) Evaluation Set
Synthetic speech clips for 101 rare drug terms, intended for evaluation only — measuring how
well an ASR system recognises rare / out-of-vocabulary medical vocabulary (target-word WER / CER /
recall). Each clip reads a real DailyMed sentence containing one target drug name, synthesised with
Google Gemini TTS across multiple voices. This is the frozen evaluation set from the master's thesis
"Audio-free lexical adaptation of Whisper's decoder"… See the full description on the dataset page: https://huggingface.co/datasets/aharalambieva/target-words-geminitts.mon-words-dataset
Dataset Summary
This dataset is a community-driven collection of the Mon language (ISO 639-3: mon). It contains 15,000+ sentences and corresponding voice recordings collected via a Mon keyboard application. The goal is to provide high-quality open-source data to support Mon language integration into global AI systems like Google Translate, OpenAI Whisper, and ChatGPT.
Supported Tasks
• Translation: Mon to English/Burmese/pali.
• ASR (Speech-to-Text): For Mon voice… See the full description on the dataset page: https://huggingface.co/datasets/Nenemin95/mon-words-dataset.
