datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
uzbek-multi-speaker-35hfante-speech-text-multispeaker_lds
Fante Speech-Text Multispeaker Dataset (LDS)
Sentence-level aligned Fante (fat) speech dataset sourced from the Church of Jesus Christ of Latter-day Saints General Conference translations.
Dataset Statistics
Split
Clips
Hours
Talks
Train
29,992
58.32
405
Eval
2,028
4.09
28
Total
32,020
62.41
433
Features
audio: 16 kHz mono FLAC sentence-level clips
text: Fante transcript (sentence-aligned)
talk_id: Source conference talk… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/fante-speech-text-multispeaker_lds.multispeaker-storycloze
Multi Speaker StoryCloze
A multispeaker spoken version of StoryCloze Synthesized with Kokoro TTS.
The dataset was synthesized to evaluate the performance of speech language models as detailed in the paper "Scaling Analysis of Interleaved Speech-Text Language Models".
We refer you to the SlamKit codebase to see how you can evaluate your SpeechLM with this dataset.
sSC and tSC
We split the generation for spoken-stroycloze and topic-storycloze as detailed in Twist.… See the full description on the dataset page: https://huggingface.co/datasets/slprl/multispeaker-storycloze.twi_multispeaker_audio_transcribed
Twi Multispeaker Audio Transcribed Dataset
Overview
The Twi Multispeaker Audio Transcribed dataset is a collection of speech recordings and their transcriptions in Asante Twi, a widely spoken dialect of the Akan language in Ghana. The dataset is designed for training and evaluating automatic speech recognition (ASR) models and other natural language processing (NLP) applications.
Dataset Details
Source: The dataset is derived from the Financial… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi_multispeaker_audio_transcribed.ga-multispeaker-speech-text-20k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Ga Multispeaker Audio Transcribed Dataset
Overview
The Ga… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ga-multispeaker-speech-text-20k.uzbek-multi-speaker-25hfante-multispeaker_speech-text-20k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Fante Multispeaker Audio Transcribed Dataset
Overview
The Fante… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/fante-multispeaker_speech-text-20k.akuapem_multispeaker_audio_transcribed
Akuapem Multispeaker Audio Transcribed Dataset
Overview
The Akuapem Multispeaker Audio Transcribed dataset is a collection of speech recordings and their transcriptions in Akuapem Twi, a widely spoken dialect of the Akan language in Ghana. The dataset is designed for training and evaluating automatic speech recognition (ASR) models and other natural language processing (NLP) applications.
Dataset Details
Source: The dataset is derived from the… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/akuapem_multispeaker_audio_transcribed.twi-speech-text-multispeaker-16k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Twi Speech-Text Parallel Dataset
Dataset Description
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-speech-text-multispeaker-16k.multispeaker-tts-ptbrDataset importado do https://gitlab.com/fb-audio-corpora
ar-eg-speech-tts-multi-speakerscml-tts-filtered-multispeaker_tokenisedfante-speech-text-multispeaker_lds
Fante Speech-Text Multispeaker Dataset (LDS)
Sentence-level aligned Fante (fat) speech dataset sourced from the Church of Jesus Christ of Latter-day Saints General Conference translations.
Dataset Statistics
Split
Clips
Hours
Talks
Train
29,992
58.32
405
Eval
2,028
4.09
28
Total
32,020
62.41
433
Features
audio: 16 kHz mono FLAC sentence-level clips
text: Fante transcript (sentence-aligned)
talk_id: Source conference talk… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/fante-speech-text-multispeaker_lds.hindi_multispeaker_datasetsomali-stt-dataset-multi-speaker-v1
Dataset Structure
The dataset contains the following columns:
text: The Somali sentence (transcription).
audio: The audio file sampled at 24,000 Hz.
speaker_id: Unique integer ID (1 to 11) representing each of the 11 speakers.
Metadata & Search Keywords
Language: Somali (so)
Speakers: 11 unique voices (balanced gender representation)
Audio Quality: 24kHz, mono, clean audio
Total Rows: 1,200
Total Duration: ~1.66 Hours (99.86 Minutes)
Intended Use: Fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/laki35/somali-stt-dataset-multi-speaker-v1.Indic-High-Fidelity-MultiSpeaker-ASR
Dataset Overview
This dataset contains high-quality multi-speaker conversational audio recordings curated for Automatic Speech Recognition (ASR) research across multiple Indic languages.
The dataset includes:
Paired audio + timestamped transcripts
Natural, non-scripted conversational speech
Dual-speaker interactions
Segment-level speaker annotations
Regionally diverse accents
Audio Specifications
Format: WAV (PCM 16-bit)
Sampling Rate: 16 kHz
Channel: Mono
Speech… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/Indic-High-Fidelity-MultiSpeaker-ASR.Synthetic-Multispeaker-Maithili-Santaliarabic-tts-saudi-multi-speaker-xtts
Arabic Saudi TTS Dataset (LJSpeech Format) 🇸🇦
This dataset is designed for training Text-to-Speech (TTS) models such as XTTS_v2 using the LJSpeech format.
📌 Overview
Language: Arabic (Saudi Dialect)
Format: LJSpeech
Use Case: TTS training (XTTS_v2, YourTTS, Tacotron, etc.)
Speakers: Multi-speaker (Male & Female)
Audio Format: WAV (mono recommended)
Sample Rate: 22050 Hz (recommended)
📂 Structure
all_data/
│
├── wavs/
│ ├── sample_0.wav
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Abdelrahman2922/arabic-tts-saudi-multi-speaker-xtts.french-b2b-tts-multispeaker
French B2B Multi-Speaker TTS Dataset
Dataset Description
A multi-speaker French text-to-speech dataset covering three B2B industry verticals: fintech/banking, e-commerce/logistics, and healthcare/medical. Audio clips are generated with diverse male and female speaker voices for conversational AI applications.
Verticals
Vertical
Description
fintech_banking
Banking operations, transfers, account inquiries, fraud alerts, investments… See the full description on the dataset page: https://huggingface.co/datasets/voxozi/french-b2b-tts-multispeaker.twi-speech-text-multispeaker-cleanmultispeakersugtts-multispeaker-max266secs-total9hrs-sr22050multispeaker-tts-sinhala\\nThis data set contains multi-speaker high quality transcribed audio data for Sinhala. The data set consists of wave files, and a TSV file.
The file si_lk.lines.txt contains a FileID, which in tern contains the UserID and the Transcription of audio in the file.
The data set has been manually quality checked, but there might still be errors.
Part of this dataset was collected by Google in Sri Lanka and the rest was contributed by Path to Nirvana organization.Novax_Multi_speakermulti_speakerspersian_multispeaker_voiceorpheus_tts_english_indian_multispeakermultispeaker-tts-ptbrDataset importado do https://gitlab.com/fb-audio-corpora
speaker_evaluation_multi_test_v0
Seamless Interaction Pairs
This dataset contains paired query and document audio clips for interaction-based
speaker evaluation. Each row describes a query clip and a related document clip,
with segment metadata and durations for analysis.
Data structure
The dataset uses a single split stored in data.parquet.
Audio files are stored under audio/ and referenced by relative paths in the
parquet file.
Columns
pair_id (string): Pair identifier.
interaction… See the full description on the dataset page: https://huggingface.co/datasets/humanify/speaker_evaluation_multi_test_v0.salt-multispeaker-eng-split
