datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
not-wake-words-speech-en
not-wake-words-speech-en
Negative (non-wake-word) speech clips, used to measure false accepts for OVOS
wake-word plugins.
Derived from the Multilingual Spoken Words Corpus
(MLCommons), which is built from Mozilla Common Voice and licensed CC-BY-4.0.
This derivative keeps the same licence and attribution requirement.
Produced with support from the NGI0 Commons Fund.
twi-words-speech-text-parallel-400k
Twi Words Speech-Text Parallel Dataset
Dataset Description
This dataset contains 413463 parallel speech-text pairs for Twi (Akan), a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Twi (Akan) - tw
Task: Speech Recognition, Text-to-Speech
Size: 413463 audio files >… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/twi-words-speech-text-parallel-400k.quran-md-words
Quran-MD - Word Level
This dataset is part of the complete Quran-MD dataset available here: Complete Quran-MD
Paper
Quran-MD: A Fine-Grained Multimodal Dataset of the Quran
(Accepted at: 5th Muslims in ML Workshop co-located with NeurIPS 2025)
📄 Paper Link: quran-md-paper
Abstract
We present Quran-MD, a comprehensive multimodal dataset of the Qur’an that integrates textual, linguistic, and audio dimensions at the verse and word levels. For each verse (ayah)… See the full description on the dataset page: https://huggingface.co/datasets/Buraaq/quran-md-words.swahili-words-speech-text-parallel
Swahili Words Speech-Text Parallel Dataset
Dataset Description
This dataset contains 411048 parallel speech-text pairs for Swahili, a widely spoken language in East Africa. The dataset consists of audio recordings paired with corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Swahili - sw
Task: Speech Recognition, Text-to-Speech
Size: 411048 audio files > 1KB… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-words-speech-text-parallel.wordshk_cantonese_speechpersian-words
Persian Words
This is a dataset of approximately 5K words, read aloud by a variety of native speakers. The dataset has been directly redistributed from this URL.
It can be used as a valuable resource for evaluating/training ASR engines or speech synthesis engines.
P.S.: I'm not the original creator of this dataset, for crediting or ownership change you can contact
all-words-in-english-with-pink-trombone
Dataset Card for Pink Trombone English Phonetic & Landmark Dataset
Repository: mcamara/all-words-in-english-with-pink-trombone
Modality: Audio + time-aligned events (landmarks) + articulatory keyframes
Language: English (IPA)
Sampling rate: 44,100 Hz (mono)
Voices: two synthetic voices — M (male) and F (female)
Summary
A large-scale, clean synthetic speech dataset generated with the Pink Trombone
articulatory synthesizer. Every English dictionary word is… See the full description on the dataset page: https://huggingface.co/datasets/mcamara/all-words-in-english-with-pink-trombone.spoken_words_en_ml_commons_filtered_splitwordshk_cantonese_speechtwi-stitched-words-asr
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
PhonemeSegmentCounting_Librispeech-wordstorgo-dysarthria-male-words-100Merged_Arabic_Corpus_of_Isolated_Words
Dataset Card for Merged_Arabic_Corpus_of_Isolated_Words
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information… See the full description on the dataset page: https://huggingface.co/datasets/arbml/Merged_Arabic_Corpus_of_Isolated_Words.PhonemeSegmentCounting_Librispeech-wordstarget-words-geminitts
Target-Word (TW) Evaluation Set
Synthetic speech clips for 101 rare drug terms, intended for evaluation only — measuring how
well an ASR system recognises rare / out-of-vocabulary medical vocabulary (target-word WER / CER /
recall). Each clip reads a real DailyMed sentence containing one target drug name, synthesised with
Google Gemini TTS across multiple voices. This is the frozen evaluation set from the master's thesis
"Audio-free lexical adaptation of Whisper's decoder"… See the full description on the dataset page: https://huggingface.co/datasets/aharalambieva/target-words-geminitts.Speech_Corpus_for_Isolated_Words
Dataset Card for Speech_Corpus_for_Isolated_Words
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/arbml/Speech_Corpus_for_Isolated_Words.50_words_datasetquran-md-words
Quran-MD - Word Level
This dataset is part of the complete Quran-MD dataset available here: Complete Quran-MD
Paper
Quran-MD: A Fine-Grained Multimodal Dataset of the Quran
(Accepted at: 5th Muslims in ML Workshop co-located with NeurIPS 2025)
📄 Paper Link: quran-md-paper
Abstract
We present Quran-MD, a comprehensive multimodal dataset of the Qur’an that integrates textual, linguistic, and audio dimensions at the verse and word levels. For each verse (ayah)… See the full description on the dataset page: https://huggingface.co/datasets/melakio/quran-md-words.eg-ADI-wordstarget_wordstarget_words_lexical_adaptation50-words-stephen-fryDataset of Stephen Fry from his singing on 50 Words For Snow by kate bush.
kash_words
