datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fante-speech-text-multispeaker_lds
Fante Speech-Text Multispeaker Dataset (LDS)
Sentence-level aligned Fante (fat) speech dataset sourced from the Church of Jesus Christ of Latter-day Saints General Conference translations.
Dataset Statistics
Split
Clips
Hours
Talks
Train
29,992
58.32
405
Eval
2,028
4.09
28
Total
32,020
62.41
433
Features
audio: 16 kHz mono FLAC sentence-level clips
text: Fante transcript (sentence-aligned)
talk_id: Source conference talk… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/fante-speech-text-multispeaker_lds.twi_multispeaker_audio_transcribed
Twi Multispeaker Audio Transcribed Dataset
Overview
The Twi Multispeaker Audio Transcribed dataset is a collection of speech recordings and their transcriptions in Asante Twi, a widely spoken dialect of the Akan language in Ghana. The dataset is designed for training and evaluating automatic speech recognition (ASR) models and other natural language processing (NLP) applications.
Dataset Details
Source: The dataset is derived from the Financial… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi_multispeaker_audio_transcribed.ga-multispeaker-speech-text-20k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Ga Multispeaker Audio Transcribed Dataset
Overview
The Ga… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ga-multispeaker-speech-text-20k.fante-multispeaker_speech-text-20k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Fante Multispeaker Audio Transcribed Dataset
Overview
The Fante… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/fante-multispeaker_speech-text-20k.akuapem_multispeaker_audio_transcribed
Akuapem Multispeaker Audio Transcribed Dataset
Overview
The Akuapem Multispeaker Audio Transcribed dataset is a collection of speech recordings and their transcriptions in Akuapem Twi, a widely spoken dialect of the Akan language in Ghana. The dataset is designed for training and evaluating automatic speech recognition (ASR) models and other natural language processing (NLP) applications.
Dataset Details
Source: The dataset is derived from the… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/akuapem_multispeaker_audio_transcribed.fante-speech-text-multispeaker_lds
Fante Speech-Text Multispeaker Dataset (LDS)
Sentence-level aligned Fante (fat) speech dataset sourced from the Church of Jesus Christ of Latter-day Saints General Conference translations.
Dataset Statistics
Split
Clips
Hours
Talks
Train
29,992
58.32
405
Eval
2,028
4.09
28
Total
32,020
62.41
433
Features
audio: 16 kHz mono FLAC sentence-level clips
text: Fante transcript (sentence-aligned)
talk_id: Source conference talk… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/fante-speech-text-multispeaker_lds.Indic-High-Fidelity-MultiSpeaker-ASR
Dataset Overview
This dataset contains high-quality multi-speaker conversational audio recordings curated for Automatic Speech Recognition (ASR) research across multiple Indic languages.
The dataset includes:
Paired audio + timestamped transcripts
Natural, non-scripted conversational speech
Dual-speaker interactions
Segment-level speaker annotations
Regionally diverse accents
Audio Specifications
Format: WAV (PCM 16-bit)
Sampling Rate: 16 kHz
Channel: Mono
Speech… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/Indic-High-Fidelity-MultiSpeaker-ASR.bengali-multi-speaker-speech-samples
Bengali Speech: Multi-Speaker Samples
This sample shows Bengali multi-speaker speech with aligned ground-truth transcripts. It is meant to help buyers review conversational structure, speaker overlap, transcript quality, and audio consistency before scoping a larger delivery.
What This Shows
Multi-speaker Bengali speech with transcript alignment
Conversation-style audio rather than isolated prompt reading
Metadata that distinguishes language, format, and speaker… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bengali-multi-speaker-speech-samples.
