datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual_librispeech
Dataset Card for MultiLingual LibriSpeech
Dataset Summary
This is a streamable version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.multilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT)
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.open-bible
OpenBibleTTS
OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license.
Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources
Source: Open Bible (CC BY-SA)
Languages
Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.multilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.open-asr-leaderboard-multilingual-datasets
ASR Leaderboard Datasets
This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS).
How to Load
To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>.
from datasets import load_dataset
# Load the FLEURS dataset for Bulgarian
fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg")
print(fleurs_bg)
# Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.multilingual-TEDX-frThe french subset of the dataset Multilingual TEDx. The data uploaded to HF corresponds to the directory fr-fr. The audio files are automatically resampled to 16 kHz.
Configs:
single_samples (default): all samples taken separately
Sample
{'file': '0u7tTptBo9I-0', 'audio': {'path': None, 'array': array([ 3.05175781e-05, 6.10351562e-05, 9.15527344e-05, ...,
-2.44140625e-04, -3.35693359e-04, -2.74658203e-04]), 'sampling_rate': 16000}, 'sentence': "Bonsoir ! Notre… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/multilingual-TEDX-fr.NEXUS-temporal_hierarchical_multi-modal
NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset
(Temporal Multimodal Slices)
This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s).
It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.MultiMed
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
ACL 2025
Khai Le-Duc, Phuc Phan, Tan-Hanh Pham, Bach Phan Tat,
Minh-Huong Ngo, Chris Ngo, Thanh Nguyen-Tang, Truong-Son Hy
Please press ⭐ button and/or cite papers if you feel helpful.
Abstract:
Multilingual automatic speech recognition (ASR) in the medical domain serves as a foundational task for various downstream applications such as speech translation, spoken language understanding, and… See the full description on the dataset page: https://huggingface.co/datasets/leduckhai/MultiMed.fante-speech-text-multispeaker_lds
Fante Speech-Text Multispeaker Dataset (LDS)
Sentence-level aligned Fante (fat) speech dataset sourced from the Church of Jesus Christ of Latter-day Saints General Conference translations.
Dataset Statistics
Split
Clips
Hours
Talks
Train
29,992
58.32
405
Eval
2,028
4.09
28
Total
32,020
62.41
433
Features
audio: 16 kHz mono FLAC sentence-level clips
text: Fante transcript (sentence-aligned)
talk_id: Source conference talk… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/fante-speech-text-multispeaker_lds.MedQA-Darija-MultiLingual
MedQA-Darija-MultiLingual
The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija.
A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region.
Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.MultiMed-ST
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
📘 EMNLP 2025
Khai Le-Duc*, Tuyen Tran*, Bach Phan Tat, Nguyen Kim Hai Bui, Quan Dang, Hung-Phong Tran, Thanh-Thuy Nguyen, Ly Nguyen, Tuan-Minh Phan, Thi Thu Phuong Tran, Chris Ngo, Nguyen X. Khanh**, Thanh Nguyen-Tang**
*Equal contribution | **Equal supervision
⭐ If you find this work useful, please consider starring the repo and… See the full description on the dataset page: https://huggingface.co/datasets/leduckhai/MultiMed-ST.Inkling-Small-Multimodal-Calibration
Inkling-Small Multimodal Calibration
The exact 1,663 samples used for BF16 routed-expert importance collection
for Inkling-Small Mixed Quant GGUF.
This is calibration material, not a held-out evaluation benchmark.
The primary balanced pass is:
Category
Samples
Valid decoder tokens
Share
Text / reasoning
462
471,858
44.976%
Code / tool-oriented source text
205
209,715
19.989%
Real image / document
486
262,476
25.018%
Real speech audio
309
105,080
10.016%
Total… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Inkling-Small-Multimodal-Calibration.multilingual-nchlt-dataset
NCHLT Auxiliary Speech Corpus - Combined Multilingual Dataset
Dataset Description
This is a combined multilingual version of the NCHLT Auxiliary Speech Corpus, compiled by the Data Science for Social Impact (DSFSI) research group at the University of Pretoria to facilitate easier benchmarking and multi-language speech recognition research.
The original auxiliary data was collected during the National Centre for Human Language Technology (NCHLT) project for the 11 official… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/multilingual-nchlt-dataset.multiturn_ks
khursanirevo/multiturn_ks
Dataset Description
Multiturn dialogue dataset with speaker-separated stereo audio and multi-language transcripts from 139 YouTube videos.
Features
Audio: Stereo audio with speaker separation (speaker 0 = left channel, speaker 1 = right channel)
Segments: Speaker turn-level annotations with timestamps for English and Malay
Multi-language: Transcripts in 9 languages (en, ms, zh-Hans, zh-Hant, ru, id, ar, ja, ko)
Video ID: YouTube video… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/multiturn_ks.multilingual-tts-benchmark
Multilingual Speech Benchmark for Zero-Shot TTS
A voice-cloning and intelligibility benchmark for six language variants, built
from Common Voice 17.0 by coverage-driven selection rather than random sampling.
Every example pairs a reference clip of one speaker with a target text that
speaker never read, so a system is asked to clone a voice and produce new
speech, which is what zero-shot TTS is actually for.
Pipeline source code:… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/multilingual-tts-benchmark.twi_multispeaker_audio_transcribed
Twi Multispeaker Audio Transcribed Dataset
Overview
The Twi Multispeaker Audio Transcribed dataset is a collection of speech recordings and their transcriptions in Asante Twi, a widely spoken dialect of the Akan language in Ghana. The dataset is designed for training and evaluating automatic speech recognition (ASR) models and other natural language processing (NLP) applications.
Dataset Details
Source: The dataset is derived from the Financial… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi_multispeaker_audio_transcribed.multivoice-synthetic-speech
Synthetic Voice Samples · Africa
Synthetic speech. No human speaker was recorded for any clip in this dataset.
Generated with afrispeech-synth: text from
africa-corpus, normalised to a
universal orthography with africa-g2p, spoken by
Google Gemini's Live API.
17,010 clips · 38.8 hours · 566 languages · 30 voices
Every clip is a distinct sentence — no sentence is repeated
Each language is read by up to 30 different voices, one sentence per voice
~1.29 hours per voice… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/multivoice-synthetic-speech.fante-multispeaker_speech-text-20k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Fante Multispeaker Audio Transcribed Dataset
Overview
The Fante… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/fante-multispeaker_speech-text-20k.ga-multispeaker-speech-text-20k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Ga Multispeaker Audio Transcribed Dataset
Overview
The Ga… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ga-multispeaker-speech-text-20k.akuapem_multispeaker_audio_transcribed
Akuapem Multispeaker Audio Transcribed Dataset
Overview
The Akuapem Multispeaker Audio Transcribed dataset is a collection of speech recordings and their transcriptions in Akuapem Twi, a widely spoken dialect of the Akan language in Ghana. The dataset is designed for training and evaluating automatic speech recognition (ASR) models and other natural language processing (NLP) applications.
Dataset Details
Source: The dataset is derived from the… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/akuapem_multispeaker_audio_transcribed.LnNor
Dataset Card for the LnNor Corpus
A multilingual dataset of high-quality speech recordings in Norwegian, English, and Polish, designed for research into cross-linguistic influence, multilingual language acquisition, and applications in NLP and speech processing such as ASR, TTS, and linguistic variability modeling. The dataset features structured experimental tasks such as reading, picture and video description, and spontaneous conversation to capture phonological, syntactic, and… See the full description on the dataset page: https://huggingface.co/datasets/MultiBridge/LnNor.arabic-multidialect
Arabic Whisper Multi-Dialect ASR Dataset
A comprehensive multi-dialect Arabic speech recognition dataset prepared for Whisper model fine-tuning.
Dataset Description
This dataset combines high-quality Arabic speech data from multiple dialects, specifically curated for fine-tuning OpenAI's Whisper models on Arabic speech recognition tasks.
Dialects Included
Modern Standard Arabic (MSA) - Formal Arabic used in media and formal contexts
Egyptian Arabic (EGY) - The… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/arabic-multidialect.MultiMed
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
ACL 2025
Khai Le-Duc, Phuc Phan, Tan-Hanh Pham, Bach Phan Tat,
Minh-Huong Ngo, Chris Ngo, Thanh Nguyen-Tang, Truong-Son Hy
Please press ⭐ button and/or cite papers if you feel helpful.
Abstract:
Multilingual automatic speech recognition (ASR) in the medical domain serves as a foundational task for various downstream applications such as speech translation, spoken language understanding, and… See the full description on the dataset page: https://huggingface.co/datasets/ttthe/MultiMed.fante-speech-text-multispeaker_lds
Fante Speech-Text Multispeaker Dataset (LDS)
Sentence-level aligned Fante (fat) speech dataset sourced from the Church of Jesus Christ of Latter-day Saints General Conference translations.
Dataset Statistics
Split
Clips
Hours
Talks
Train
29,992
58.32
405
Eval
2,028
4.09
28
Total
32,020
62.41
433
Features
audio: 16 kHz mono FLAC sentence-level clips
text: Fante transcript (sentence-aligned)
talk_id: Source conference talk… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/fante-speech-text-multispeaker_lds.multilingual_librispeech_french_phoneme
Multilingual LibriSpeech French Phoneme
Dataset Summary
This dataset is a curated version of the French subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme).
The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into French acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_french_phoneme.arabic-multidialect-emotional-speech-demo
DataHive AI — Demo: Arabic Multi-Dialect Emotional Speech
A DataHive AI dataset: a stratified 1-hour demo sample from a full corpus of 50+ hours. We can also create larger audio datasets upon client request.
Most public Arabic speech corpora flatten dialect into a single label and ignore emotion entirely. This corpus does the opposite: every recording is tagged with one of four regional Arabic dialects (Najdi, Hejazi, Jordanian, Moroccan) and one of four target emotions (Sad, Happy… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/arabic-multidialect-emotional-speech-demo.Indic-High-Fidelity-MultiSpeaker-ASR
Dataset Overview
This dataset contains high-quality multi-speaker conversational audio recordings curated for Automatic Speech Recognition (ASR) research across multiple Indic languages.
The dataset includes:
Paired audio + timestamped transcripts
Natural, non-scripted conversational speech
Dual-speaker interactions
Segment-level speaker annotations
Regionally diverse accents
Audio Specifications
Format: WAV (PCM 16-bit)
Sampling Rate: 16 kHz
Channel: Mono
Speech… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/Indic-High-Fidelity-MultiSpeaker-ASR.panta_instruct_multi_modal_v1
Panta Instruct Multi-Modal v1
Dataset d'instructions multimodal en français : chaque exemple associe une question
(texte + parole + pictogrammes) à une réponse (texte + pictogrammes).
Colonnes
Colonne
Type
Description
audio
Audio (24 kHz, mono)
Enregistrement de la question (text_input)
text_input
string
Question / instruction
text_output
string
Réponse
pictos_input
list[string]
Identifiants des pictogrammes de la question
pictos_output… See the full description on the dataset page: https://huggingface.co/datasets/audibeal74/panta_instruct_multi_modal_v1.multilingual-synthetic-tts
Multilingual Synthetic TTS Dataset
🏆 Submitted to the Uncharted Data Challenge
hosted by Adaption Labs — credit to
Adaptive Data by Adaption for organizing the hackathon.
A large-scale synthetic multilingual speech dataset — 68,677 clips across
9 languages, generated with Qwen3-TTS-12Hz-1.7B-Base
using zero-shot voice cloning from 5 reference speakers.
Intended for training and evaluating TTS, ASR, voice conversion, and
multilingual speech models. Each clip is paired with the… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-synthetic-tts.multilingual_librispeech_fr_punctuated
Multilingual LibriSpeech French (Punctuated)
This dataset is a converted version of BrunoHays/multilingual_librispeech_fr_punctuated
in the new Hugging Face datasets format (Parquet-based, without loading scripts).
Original Dataset
The original dataset contains French speech data from Multilingual LibriSpeech with punctuated transcriptions.
Changes
Converted from old loading script format to new Parquet-based format
Maintains all original features and data… See the full description on the dataset page: https://huggingface.co/datasets/antoineedy/multilingual_librispeech_fr_punctuated.
