datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twi-trigrams-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-trigrams-speech-text-parallel.makhuwa-trigrams-speech-text-parallel
Makhuwa Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 154253 parallel speech-text pairs for Makhuwa, a language spoken primarily in Mozambique. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Makhuwa - vmw
Task: Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/makhuwa-trigrams-speech-text-parallel.twi-trigrams-speech-text-parallel
Twi Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Twi - twi
Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/twi-trigrams-speech-text-parallel.chichewa-trigrams-speech-text-parallel
Chichewa Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 132549 parallel speech-text pairs for Chichewa, a language spoken primarily in Malawi. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Chichewa - ny
Task: Speech Recognition… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/chichewa-trigrams-speech-text-parallel.cm.trial
Dataset Card for Common Voice Corpus 11.0
Dataset Summary
The Common Voice dataset consists of a unique MP3 and corresponding text file.
Many of the 24210 recorded hours in the dataset also include demographic metadata like age, sex, and accent
that can help improve the accuracy of speech recognition engines.
The dataset currently consists of 16413 validated hours in 100 languages, but more voices and languages are always added.
Take a look at the Languages page to… See the full description on the dataset page: https://huggingface.co/datasets/taqwa92/cm.trial.TRILOGUE
TRILOGUE
TRILOGUE is a trilingual spoken-dialogue fact-checking benchmark with clean
text, turn-level ASR transcripts, word-level timestamp alignments, evidence
supervision, fixed article-disjoint experimental splits, and paired audio. The
complete benchmark contains 11,957 dialogues, 187,544 turns, and 390.3 hours of
audio in English, Russian, and Kazakh. The current package releases 8,036
Russian and Kazakh recordings (242.5 hours), including 4,998 human-read
recordings; 3,921… See the full description on the dataset page: https://huggingface.co/datasets/chaewanC/TRILOGUE.kinyarwanda-speech-trimmed
Kinyarwanda Speech Dataset (Trimmed)
This dataset contains processed Kinyarwanda speech data with trimmed audio segments.
Dataset Structure
The dataset contains two splits:
dev_test: 9,263 samples
test: 9,265 samples
Features
Each sample contains:
id: Unique identifier for the sample
audio: Audio data
audio_language: Language of the audio (Kinyarwanda)
text: Transcription of the audio
prompt: Associated prompt or context
duration: Duration of the audio… See the full description on the dataset page: https://huggingface.co/datasets/yigagilbert/kinyarwanda-speech-trimmed.marathi-tts-yt-trimmed
Marathi TTS YouTube Trimmed Audio Corpus
Dataset Summary
Marathi TTS YouTube Trimmed Audio Corpus comprises preprocessed, segmented, and silence-trimmed Marathi speech audio clips derived from curated public Marathi educational, monologue, and conversational videos. Each clip is carefully aligned with clean Marathi text transcriptions for high-clarity speech model training.
Dataset Structure
Format: Segmented audio files (WAV, 22.05 kHz) + text… See the full description on the dataset page: https://huggingface.co/datasets/Srijan-Upadhyay/marathi-tts-yt-trimmed.
