datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual_librispeech
Dataset Card for MultiLingual LibriSpeech
Dataset Summary
This is a streamable version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.multilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT)
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.multilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.open-bible
OpenBibleTTS
OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license.
Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources
Source: Open Bible (CC BY-SA)
Languages
Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.multilingual-TEDX-frThe french subset of the dataset Multilingual TEDx. The data uploaded to HF corresponds to the directory fr-fr. The audio files are automatically resampled to 16 kHz.
Configs:
single_samples (default): all samples taken separately
Sample
{'file': '0u7tTptBo9I-0', 'audio': {'path': None, 'array': array([ 3.05175781e-05, 6.10351562e-05, 9.15527344e-05, ...,
-2.44140625e-04, -3.35693359e-04, -2.74658203e-04]), 'sampling_rate': 16000}, 'sentence': "Bonsoir ! Notre… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/multilingual-TEDX-fr.open-asr-leaderboard-multilingual-datasets
ASR Leaderboard Datasets
This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS).
How to Load
To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>.
from datasets import load_dataset
# Load the FLEURS dataset for Bulgarian
fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg")
print(fleurs_bg)
# Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.Multilingual-TTS-language
Multilingual-TTS-language
malaysia-ai/Multilingual-TTS
with two extra columns:
column
description
audio_filename, text, speaker
unchanged from malaysia-ai/Multilingual-TTS
language
language detected from the text (transcript) column of every row
post-normalized
text after rule-based punctuation / capitalization normalization
All original columns and the file/folder layout are preserved: 1552 subsets /
1561 parquet files, 121,842,371 rows (August 2026). Every… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS-language.multilingual-accent-speech
🎙️ Silencio Network: Voice AI Sample Dataset
📊 This is a sample. The full Silencio corpus contains 100,000+ hours across 170+ countries and 100+ languages.
📧 Contact: sofia@silencioai.com for custom datasets, bulk licensing, or specific language requests.
🌍 Why Silencio Data?
Silencio data is collected in the wild from a massive, opt-in community (2M+ contributors across 180+ countries), giving you:
✅ Real-world accents, dialects, devices, and… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/multilingual-accent-speech.MedQA-Darija-MultiLingual
MedQA-Darija-MultiLingual
The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija.
A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region.
Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.multilingual-nchlt-dataset
NCHLT Auxiliary Speech Corpus - Combined Multilingual Dataset
Dataset Description
This is a combined multilingual version of the NCHLT Auxiliary Speech Corpus, compiled by the Data Science for Social Impact (DSFSI) research group at the University of Pretoria to facilitate easier benchmarking and multi-language speech recognition research.
The original auxiliary data was collected during the National Centre for Human Language Technology (NCHLT) project for the 11 official… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/multilingual-nchlt-dataset.multilingual-tts-benchmark
Multilingual Speech Benchmark for Zero-Shot TTS
A voice-cloning and intelligibility benchmark for six language variants, built
from Common Voice 17.0 by coverage-driven selection rather than random sampling.
Every example pairs a reference clip of one speaker with a target text that
speaker never read, so a system is asked to clone a voice and produce new
speech, which is what zero-shot TTS is actually for.
Pipeline source code:… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/multilingual-tts-benchmark.multilingual_librispeechMultilingual LibriSpeech (MLS) dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.Zambia-MultiLingual-ASR-Dataset
🇿🇲 Zambia Multilingual ASR Dataset
A continuously growing and curated multilingual speech corpus for Zambian languages, designed to advance Automatic Speech Recognition (ASR) research through community-driven data collection and real-world evaluation.
Overview
The Zambia Multilingual ASR Dataset is an open, continuously evolving speech corpus developed as part of the ZamVoice project.
The dataset supports research and development of Automatic Speech… See the full description on the dataset page: https://huggingface.co/datasets/buumba641/Zambia-MultiLingual-ASR-Dataset.multilingual_librispeech_french_phoneme
Multilingual LibriSpeech French Phoneme
Dataset Summary
This dataset is a curated version of the French subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme).
The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into French acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_french_phoneme.synthetic-multilingual-speaker-diarization
Synthetic Multilingual Speaker Diarization Dataset
This dataset contains synthetic multilingual speaker diarization data with Hindi, English, and Punjabi audio samples.
Dataset Structure
├── audio/ # WAV audio files (16kHz) - 3417 files
├── all_samples_combined.csv # Complete dataset annotations (with silence)
└── all_visible_combined.csv # Visible dataset annotations (without silence)
Statistics
Total samples: 3417 audio… See the full description on the dataset page: https://huggingface.co/datasets/noty7gian/synthetic-multilingual-speaker-diarization.multilingual-synthetic-tts
Multilingual Synthetic TTS Dataset
🏆 Submitted to the Uncharted Data Challenge
hosted by Adaption Labs — credit to
Adaptive Data by Adaption for organizing the hackathon.
A large-scale synthetic multilingual speech dataset — 68,677 clips across
9 languages, generated with Qwen3-TTS-12Hz-1.7B-Base
using zero-shot voice cloning from 5 reference speakers.
Intended for training and evaluating TTS, ASR, voice conversion, and
multilingual speech models. Each clip is paired with the… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-synthetic-tts.Multilingual_Speech_Dataset
Multilingual Speech Dataset
Paper: A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English
Repository: https://github.com/IS2AI/MultilingualASR
Description: This repository provides the dataset used in the paper "A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English". The paper focuses on training a single end-to-end (E2E) ASR model for Kazakh, Russian, and English, comparing monolingual and multilingual approaches… See the full description on the dataset page: https://huggingface.co/datasets/issai/Multilingual_Speech_Dataset.multilingual_librispeech_spanish_phoneme
Multilingual LibriSpeech Spanish Phoneme
Dataset Summary
This dataset is a curated version of the Spanish subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme).
The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into Spanish acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_spanish_phoneme.Adaption-multilingual-speech
This dataset is a remastered version of
Reubencf/multilingual-synthetic-tts
prepared using Adaption's Adaptive Data platform.
Multilingual Speech (Adaption)
10,274 audio + text rows selected from the original 68,677-clip
multilingual synthetic speech corpus, with Adaption-sharpened
enhanced_prompt and enhanced_completion columns. Every row carries
the synthesised audio, the ground-truth text, and language/style/voice
metadata — ready for speech SFT.
Original dataset (for… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-multilingual-speech.multilingual-speech
Multilingual Indian Conversational Speech
A dataset of naturalistic, spontaneous two-speaker conversations across
13 Indian languages, with segment-level transcripts, speaker profiles,
timestamps, and recording metadata. Designed for ASR, TTS, speaker
diarization, and conversational speech research.
Languages (13)
Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Nepali,
Odia, Punjabi, Tamil, Telugu, Urdu.
Content
Conversations… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/multilingual-speech.leyu-ethiopian-multilingual-speech-corpus-2026
🇪🇹 Leyu Ethiopian Multilingual Speech Corpus 2026
Official research-grade multilingual speech dataset submitted for the Leyu Platform Data Collection Competition 2026, organized by gheero (Leyu Platform Team).
🏢 Platform & Team Verification Details
Team Name: SoundWaveET
Hugging Face Organization: SoundWaveET
Dataset Repository: SoundWaveET/leyu-ethiopian-multilingual-speech-corpus-2026
Platform Live Frontend: https://leyusound.netlify.app
Platform Live API:… See the full description on the dataset page: https://huggingface.co/datasets/SoundWaveET/leyu-ethiopian-multilingual-speech-corpus-2026.multilingual-accent-speech-v2
Silencio Network: Multilingual Accent Speech Dataset (Sample)
Overview
Silencio data is valuable because it's collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don't capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed)… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/multilingual-accent-speech-v2.multilingual_librispeech_fr_punctuated
Multilingual LibriSpeech French (Punctuated)
This dataset is a converted version of BrunoHays/multilingual_librispeech_fr_punctuated
in the new Hugging Face datasets format (Parquet-based, without loading scripts).
Original Dataset
The original dataset contains French speech data from Multilingual LibriSpeech with punctuated transcriptions.
Changes
Converted from old loading script format to new Parquet-based format
Maintains all original features and data… See the full description on the dataset page: https://huggingface.co/datasets/antoineedy/multilingual_librispeech_fr_punctuated.sound-wave-leyu-ethiopian-multilingual-speech-corpus-2026
🇪🇹 Leyu Ethiopian Multilingual Speech Corpus 2026
Official research-grade multilingual speech dataset submitted for the Leyu Platform Data Collection Competition 2026, organized by gheero (Leyu Platform Team).
🏢 Platform & Team Verification Details
Team Name: SoundWaveET
Hugging Face Organization: SoundWaveET
Dataset Repository: SoundWaveET/leyu-ethiopian-multilingual-speech-corpus-2026
Platform Live Frontend: https://leyusound.netlify.app
Platform Live API:… See the full description on the dataset page: https://huggingface.co/datasets/LeyuCompetition/sound-wave-leyu-ethiopian-multilingual-speech-corpus-2026.multilingual_librispeech_italian_phoneme
Multilingual LibriSpeech Italian Phoneme
Dataset Summary
This dataset is a curated version of the Italian subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme).
The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into Italian acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_italian_phoneme.multilingual-tts-corpus
Multilingual TTS Corpus
A multilingual text-to-speech dataset containing audio recordings with text transcriptions across multiple languages. This is the first batch; more languages will be added over time.
Languages (Batch 1)
Subset
Language
Audio Files
Annotation Format
Source
ru-tts/
Russian
10
JSONL (single file)
俄语TTS(文本标注及语音)基石数据集 #294
ru-speech/
Russian
5
JSONL (single file)
俄语高质量语音音频语料库 #481
vi-speech/
Vietnamese
5
JSONL (single file)
越南语高质量语音音频语料库… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/multilingual-tts-corpus.mms-multilingual-audio-5to30min
Multilingual Audio Dataset (5-30min)
This dataset contains continuous speech audio files for various languages (ranging from 5 to 30 minutes in length per language) collected from diverse sources including Hugging Face and YouTube.
Dataset Statistics
Total Languages: 100
Sources: HF Omnilingual ASR Corpus, YouTube
Language Details
Language Code
Language Name
Source
Duration (seconds)
jpn
Japanese
youtube
1459.84
eng
English
youtube… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/mms-multilingual-audio-5to30min.alconost-multilingual-speech-gold
Multilingual Speech & Translation Dataset — EN↔JA/AR-EG/PL/RU (10 phrases, dual-take)
Description
10 English source phrases with expert human translations into Japanese, Egyptian
Arabic (ar-EG), and Polish. Each target phrase is recorded by native speakers (two
takes each). Audio files are WAV 48 kHz mono, 16‑bit PCM format.
Translations are produced and QA'd by professional linguists; recordings follow
consistent orthography/style (AR-EG: Egyptian dialect; JA/PL: standard). All… See the full description on the dataset page: https://huggingface.co/datasets/alconost/alconost-multilingual-speech-gold.open-asr-leaderboard-multilingual-datasets
Open ASR Leaderboard Armenian Test Datasets
This private repository holds leaderboard-compatible Armenian test
configurations while their integration is being validated.
Configurations
fleurs_hy
Source: google/fleurs,
configuration hy_am, test split
Reviewed reference changes:
Metric-AI/fleurs-corrections,
test split
932 recordings; all 314 reviewed corrections were matched to the original
source transcript and applied
mcv_hy… See the full description on the dataset page: https://huggingface.co/datasets/Metric-AI/open-asr-leaderboard-multilingual-datasets.multilingual-asr
multilingual-asr
CoVoST 2 and Common Voice 17.0 Swahili and Hausa, re-packaged under one feature schema so the
configs can be concatenated into a single multi-task training mix. Four configs, ~206 hours of
distinct audio, 8.44 GB of Parquet. No Bambara.
Load
from datasets import load_dataset
asr = load_dataset("djelia/multilingual-asr", "covost2-transcription", split="train")
sw_test = load_dataset("djelia/multilingual-asr", "swahili", split="test")… See the full description on the dataset page: https://huggingface.co/datasets/djelia/multilingual-asr.
