datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ParsVoice
ParsVoice
A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
📣 Accepted to the EMNLP 2026 Main Conference.
ParsVoice is the largest publicly available Persian speech–text corpus tailored for
training multi-speaker text-to-speech (TTS) systems. It is built from long-form
Persian audiobook recordings using a fully automated pipeline combining sentence-aware
segmentation, ASR transcription, a ParsBERT sentence-completion classifier, binary-search… See the full description on the dataset page: https://huggingface.co/datasets/MohammadJRanjbar/ParsVoice.Bagpiper_PreTrain_Data
Bagpiper Pretraining Data
Bagpiper Pretraining Data is the public rich-captioned audio snapshot associated
with Bagpiper, an open-ended audio language
model that learns bidirectional mappings between audio and comprehensive text
descriptions across speech, music, environmental sound, and mixtures.
The en metadata describes the primary rich-caption language. Source audio can
contain speech or singing in other languages; it is not an English-only audio
guarantee.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_PreTrain_Data.common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.quran-tajweed-phonetics
The complete phonetic layer of the Quran in the riwaya of Hafs 'an
'Asim via tariq al-Shatibiyyah: 6,236 ayat, 522,475 phones, every
phone carrying its tajweed attribution: madd class with its transmitted
length range, ghunna grade, qalqalah class, tafkheem with its rank, sakt,
the seventeen sifat, and the rule that produced it.
Built and maintained by Quran Lab, a waqf building open technology in
the service of the Quran.
How it was built and verified
Indexed from the… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quran-tajweed-phonetics.mls-eng-speaker-descriptions
Dataset Card for Annotations of English MLS
This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-speaker-descriptions.bangla-10k
Bangla-10K: A Challenging, Metadata-Rich Corpus of Read and Conversational Bengali Speech from India and Bangladesh
Bangla-10K is a 10,816-hour Bengali speech corpus with
624,951 recordings from India and Bangladesh: a 10,070.8-hour core corpus
(567,323 recordings) and a separately collected 745.1-hour evaluation set
(57,628 recordings). It combines scripted single-speaker read speech with
natural multi-speaker conversations for Bengali automatic speech recognition
(ASR).
The… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bangla-10k.chinese-lips-speech-slide-probe
Chinese-LiPS Speech + Slide Probe
A self-contained probe set for testing whether visual slide context helps
simultaneous speech translation — with the input as audio, not transcripts.
Why audio matters: feeding a transcript to a text LLM deletes the acoustic
ambiguity (homophones, polysemy) that slide context is meant to resolve; the
transcript already commits to one reading. Any honest test of "does vision help
streaming ST" must consume speech.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-speech-slide-probe.pao-audio-dataset
🎙️ Pa'O Audio Dataset
ပအိုဝ်ႏ အငေါဝ်း အဆင်ႏဗာႏ ရွမ်ခြွဉ်းဗူႏ
📌 Project Summary
The Pa'O Audio Dataset is an open-source initiative created to facilitate the development of speech technologies and Artificial Intelligence tools for the Pa'O language (ပအိုဝ်ႏဘာႏသာႏငေါဝ်းငွါ).
Pa'O is primarily spoken in Shan State and other regions of Myanmar. As a low-resource language in the AI landscape, this dataset provides audio recordings and corresponding… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-audio-dataset.pavo-bench
PAVO-Bench: 50K-Turn Benchmark for ASR-LLM-TTS Pipeline Routing
Code: github.com/vnmoorthy/pavo-bench · Paper: TMLR 2026 (accepted) · Authors: NarasingaMoorthy VeiluKanthaPerumal (UPenn), Mohammed Imthathullah (Google)
pip install git+https://github.com/vnmoorthy/pavo-bench.git
Headline results (vs fixed-cloud baseline, 50,000 voice turns)
Metric
Result
Significance
P95 end-to-end latency (H100, LibriSpeech)
−10.3% (−167 ms)
—
Median latency
−34%… See the full description on the dataset page: https://huggingface.co/datasets/vnmoorthy/pavo-bench.mls-eng-10k-tags_tagged_10k_generated
Dataset Card for Annotations of 10K hours of English MLS
This dataset consists in annotations of a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-10k-tags_tagged_10k_generated.korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.knesset-plenums
About
This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) plenums as part of the ivrit.ai project.
Consider visiting the preview space for this dataset here
Method
Data dumps from the Knesset contain A/V recordings, alongside proprietary protocols with timestamps.
We extract the audio stream, and clean up timestamp mistakes (such as backward jumps, or out-of-order timestamp artifacts).
The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums.marathi-phonology-matrices
मराठी व्याकरण आणि ध्वनी मॅट्रिक्स
Marathi Phonology Matrices
गणितीय ध्वनी संश्लेषणासाठी (Mathematical Speech Synthesis) तयार केलेला सर्वसमावेशक मराठी फोनोलॉजी डेटासेट.
🎯 उद्देश्य
हा डेटासेट मराठी भाषेच्या:
फोनोलॉजिकल विश्लेषण
मॉर्फोलॉजी (लिंग, वचन, काळ)
संधि व श्व नियम
युक्तक्षर (Clusters)
Duration & Pitch नियम
Loanword adaptation
या सर्वांसाठी संरचित डेटा पुरवतो. TTS, ASR, G2P आणि Computational Linguistics संशोधनासाठी उपयुक्त.
📊… See the full description on the dataset page: https://huggingface.co/datasets/kalpesh77/marathi-phonology-matrices.polish-tedx-asr-eval
Polish-TEDx-ASR-Eval
A dataset for evaluating automatic speech recognition (ASR) systems for Polish in the domain of TEDx public talks.
Contains audio segments from Polish TEDx talks available on YouTube (CC BY-NC-ND 4.0) and synthetic speech generated with KugelAudio (MIT), with manually created and cross-verified transcriptions. Created as part of the course "Workshops on Evaluation of Speech Recognition Systems" (ZWESUI, AMU 2026) by Group 1.
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/s512757/polish-tedx-asr-eval.preprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2605
5bfe2d098c8486d97fac8be76d86ec9146435245
train
56:46:32
50,557
589,095
11.7
31.9
techiaith/corpws-clllc-wlga
5d00294c31c78b1d7937bb2c2bc6cc70bc18d410
clips
48:20:49
27,579… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607.paralingua_ru
Russian Paralinguistic Annotation Dataset
Датасет паралингвистической разметки спикеров из трёх русскоязычных корпусов:
biggest_ru_book,
DeepSpeech и Golos.
Что размечалось
Каждое аудио размечалось вручную по следующим характеристикам:
Поле
Описание
Пример значений
gender
Пол спикера
мужской, женский
age_group
Возрастная группа
молодой, взрослый, пожилой
voice_pitch
Высота голоса
низкий, средний, высокий
loudness
Громкость
тихий, нормальный… See the full description on the dataset page: https://huggingface.co/datasets/turnipseason/paralingua_ru.pinga-fogo-chico-xavier
🎙️ Pinga-Fogo com Chico Xavier — TV Tupi, 1971
As duas entrevistas históricas do médium Chico Xavier, transmitidas ao vivo pela
TV Tupi em 1971, transcritas e estruturadas em turnos de fala com timestamp.
345 turnos (115 deles respostas do próprio Chico Xavier), a partir de
6 horas de áudio — o registro mais extenso do médium falando de improviso,
sem edição, diante de um painel de jornalistas.
Arquivos
Arquivo
Programa
Turnos
Respostas do Chico… See the full description on the dataset page: https://huggingface.co/datasets/ia-espirita/pinga-fogo-chico-xavier.persian-asr-text-2.69M-deduped
🗂️ persian-asr-text-2.69M-deduped
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Deduplicated Persian ASR text dataset used by the training stack.
پیکرهٔ متنی فارسیِ حذفتکرارشده برای ساخت واژگان، مدلسازی زبانی و پشتیبانی از آموزش ASR.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
4 files; approximately 109.64 MB
4 فایل؛ حدود 109.64… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-text-2.69M-deduped.tajik-asr-youtube
tajik-asr-youtube
Every Tajik-language clip from this project's YouTube scrape — 41 channels of news, talk
shows, podcasts, audiobooks, and learning content — with machine transcripts and the
verification scores left in as columns instead of applied as a filter. Pick your own
quality threshold; the training corpus this project actually ships
(tajik-asr-corpus-v3)
is the gated subset.
Layout
Parquet shards under data/ with an audio struct column (16 kHz mono FLAC… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-youtube.pseudolabel-malaya-speech-stt-train-whisper-large-v3mls-eng-10k-tags_tagged_10k_generated
Dataset Card for Annotations of 10K hours of English MLS
This dataset consists in annotations of a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours… See the full description on the dataset page: https://huggingface.co/datasets/pharaouk/mls-eng-10k-tags_tagged_10k_generated.uk-pods
uk-pods - speech datasets of Ukrainian podcasts.
Preparation
Clone the dataset repository and extract the content of clips.tar.gz archive.
git clone https://huggingface.co/datasets/taras-sereda/uk-pods
cd uk-pods && tar -zxvf clips.tar.gz
To use these manifests for training/inference with NeMo [1] modify audio_filepath to absolute locations of audio files extracted in previous step.
# data_root=<clonned_repo_dir> # /home/ubuntu/uk-pods
data_root=$(realpath .)
sed -i… See the full description on the dataset page: https://huggingface.co/datasets/taras-sereda/uk-pods.ftspeech-pnc-da
Dataset Card for ftspeech-pnc-da
Dataset Summary
RyeAI/ftspeech-pnc-da is a text-only Danish punctuation and capitalization companion dataset derived from the original Hugging Face dataset alexandrainst/ftspeech: https://huggingface.co/datasets/alexandrainst/ftspeech
Each row contains restored punctuated text for an existing FTSpeech training utterance together with identifiers that allow the row to be joined back to the original source dataset:
utterance_id:… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/ftspeech-pnc-da.zamai-pashto-voice2voice
ZamAI Pashto Voice2Voice
This dataset contains Pashto voice-to-voice preparation metadata for speech and translation experiments. It focuses on Pashto speech records, dialect information, transcript text, and a small viewer-ready sample manifest.
Configs
from datasets import load_dataset
metadata = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "metadata")
sample = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "viewer_sample")
Files… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-voice2voice.PapaChaves
PapaChaves 🇨🇷
PapaChaves is a longitudinal corpus of presidential press conferences from the administration of Rodrigo Chaves Robles, President of Costa Rica (May 2022 – May 2026). The dataset contains automatic speech transcriptions of 308 press conferences, covering the full presidential term.
"PapaChaves" was a nickname given to President Chaves that leaked into the press during his administration.
Dataset Summary
Stat
Value
Videos
308
Total audio… See the full description on the dataset page: https://huggingface.co/datasets/ergar/PapaChaves.librispeech_asr
Dataset Card for librispeech_asr
LibriSpeech ASR 2s Splits Dataset
Version of LibriSpeech ASR corpus split into 2s clips.
Usage
from datasets import load_dataset
# Load the dataset from the Hub
dataset = load_dataset("pavanyellow/librispeech_asr")
# Or load a specific split
dataset = load_dataset("pavanyellow/librispeech_asr", split="train")
# Access the data
for example in dataset['train'][:5]:
audio = example['audio']
text = example['text']
zamai-pashto-voice2voice
ZamAI Pashto Voice2Voice
Languages: psLicense: cc-by-4.0Task categories: automatic-speech-recognition, audio-to-audioSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for automatic-speech-recognition, audio-to-audio tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/zamai-pashto-voice2voice")
print(dataset)
Configs
default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-voice2voice.coral-v3-conversation-pnc-da
CoRal v3 Conversation PnC DA
RyeAI/coral-v3-conversation-pnc-da is a text-only Danish punctuation and
capitalization companion for the conversation training split of
CoRal-project/coral-v3.
It contains 102,226 restored transcript rows and no audio bytes.
Pinned companion revision: 6e4fbafde87fbffadd58bbe39a3a2e09e884351a.
SHA-256 of data/train-00000-of-00001.parquet:
79d30671f238e884a5b71b682bc811043e7df075566ce5565a236e66ffe32000.
Each row can be joined back to the gated source… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/coral-v3-conversation-pnc-da.pi_bench
pi-bench
pi-bench is a multi-task audio benchmark prepared for public hosting and evaluation reproducibility. The repository is organized as a Hugging Face dataset with one dataset config per task file under data/, so each benchmark subset is visible and loadable independently.
Overview
The current release contains 11 task-specific configs spanning three broad categories:
Counterfactual/contextual QA (CTC_*)
Clarification-seeking QA (Trivia_Clarification_*… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-pi-bench/pi_bench.
