multi-speaker
uzbek-multi-speaker-35hfante-speech-text-multispeaker_lds
Fante Speech-Text Multispeaker Dataset (LDS)
Sentence-level aligned Fante (fat) speech dataset sourced from the Church of Jesus Christ of Latter-day Saints General Conference translations.
Dataset Statistics
Split
Clips
Hours
Talks
Train
29,992
58.32
405
Eval
2,028
4.09
28
Total
32,020
62.41
433
Features
audio: 16 kHz mono FLAC sentence-level clips
text: Fante transcript (sentence-aligned)
talk_id: Source conference talk… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/fante-speech-text-multispeaker_lds.multispeaker-storycloze
Multi Speaker StoryCloze
A multispeaker spoken version of StoryCloze Synthesized with Kokoro TTS.
The dataset was synthesized to evaluate the performance of speech language models as detailed in the paper "Scaling Analysis of Interleaved Speech-Text Language Models".
We refer you to the SlamKit codebase to see how you can evaluate your SpeechLM with this dataset.
sSC and tSC
We split the generation for spoken-stroycloze and topic-storycloze as detailed in Twist.… See the full description on the dataset page: https://huggingface.co/datasets/slprl/multispeaker-storycloze.twi_multispeaker_audio_transcribed
Twi Multispeaker Audio Transcribed Dataset
Overview
The Twi Multispeaker Audio Transcribed dataset is a collection of speech recordings and their transcriptions in Asante Twi, a widely spoken dialect of the Akan language in Ghana. The dataset is designed for training and evaluating automatic speech recognition (ASR) models and other natural language processing (NLP) applications.
Dataset Details
Source: The dataset is derived from the Financial… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi_multispeaker_audio_transcribed.ga-multispeaker-speech-text-20k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Ga Multispeaker Audio Transcribed Dataset
Overview
The Ga… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ga-multispeaker-speech-text-20k.uzbek-multi-speaker-25h
