datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
genshin-voice-v3.3-mandarin
Dataset Card for Genshin Voice
Dataset Description
Dataset Summary
The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game.
Languages
The text in the dataset is in Mandarin.
Dataset Creation
Source Data
Initial Data Collection and Normalization
The data was obtained by unpacking the Genshin Impact game.
Who are the source language producers?
The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.knesset-committees
About
This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) committee sessions as part of the ivrit.ai project.
Consider visiting the preview space for this dataset here
Method
Data dumps from the Knesset contain A/V recordings of committee sessions, alongside human-generated protocols.
We extract the audio stream, abd produce weakly time stamped segmentation of the protocol text (we… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-committees.AISHELL-1Aishell is an open-source Chinese Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd.
400 people from different accent areas in China are invited to participate in the recording, which is conducted in a quiet indoor environment using high fidelity microphone and downsampled to 16kHz. The manual transcription accuracy is above 95%, through professional speech annotation and strict quality inspection. The data is free for academic use. We hope to provide moderate amount… See the full description on the dataset page: https://huggingface.co/datasets/AISHELL/AISHELL-1.everyayah﷽
Dataset Card for Tarteel AI's EveryAyah Dataset
Dataset Summary
This dataset is a collection of Quranic verses and their transcriptions, with diacritization, by different reciters.
Supported Tasks and Leaderboards
[Needs More Information]
Languages
The audio is in Arabic.
Dataset Structure
Data Instances
A typical data point comprises the audio file audio, and its transcription called text.
The duration… See the full description on the dataset page: https://huggingface.co/datasets/tarteel-ai/everyayah.ZeroSpeech
ZeroSpeech
A large synthetic Vietnamese speech corpus for ASR training: 9,867,987
utterances / 26,896 hours, spoken by 199,265
distinct voices, generated with
ZeroTTS from web and
conversational text.
Every clip is 16 kHz mono FLAC, 1–30 s, paired with the exact text it was
synthesized from.
Fields
field
type
description
audio
Audio(16 kHz)
the waveform, FLAC-encoded
text
string
the transcript — the exact string given to the TTS
source
string
which… See the full description on the dataset page: https://huggingface.co/datasets/zeroweight-ai/ZeroSpeech.audio-v2This dataset contains >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license.
It wa released on April 20th, 2025.
You can find the full list of sources in this dataset under the dataset's sources.txt.
Paper: https://arxiv.org/abs/2307.08720
If you use our datasets, the following quote is preferable:
@misc{marmor2023ivritai,
title={ivrit.ai: A Comprehensive Dataset of Hebrew Speech for AI Research and Development},
author={Yanir Marmor and Kinneret Misgav and Yair… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2.synthetic_dem
Dataset Card for synthetic_dem
Dataset Summary
The Synthetic DEM Corpus is the result of the first phase of a collaboration between El Colegio de México (COLMEX) and the Barcelona Supercomputing Center (BSC).
It all began when COLMEX was looking for a way to have its Diccionario del Español de México (DEM), which can be accessed online, include the option to play each of its words with a Mexican accent through synthetic speech files. On the other hand, BSC is always on… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/synthetic_dem.tarteel-ai-everyayah-Quran﷽
Dataset Card for Tarteel AI's EveryAyah Dataset
Dataset Summary
This dataset is a collection of Quranic verses and their transcriptions, with diacritization, by different reciters.
How to download
!pip install -q datasets
from datasets import load_dataset
dataset =load_dataset("Salama1429/tarteel-ai-everyayah-Quran", verification_mode="no_checks")
Supported Tasks and Leaderboards
[Needs More Information]
Languages
The audio is in… See the full description on the dataset page: https://huggingface.co/datasets/Salama1429/tarteel-ai-everyayah-Quran.Neyshekar
Neyshekar
Neyshekar is an open, community-driven Persian speech dataset collected via a web-based crowdsourcing platform at https://ney.shekar.io. It is designed to support research and development in text-to-speech (TTS), automatic speech recognition (ASR), speech representation learning, and other downstream Persian speech applications.
The recordings are provided by volunteer contributors, all of whom are native Persian speakers. Each release represents a… See the full description on the dataset page: https://huggingface.co/datasets/shekar-ai/Neyshekar.Multilingual-TTS-language
Multilingual-TTS-language
malaysia-ai/Multilingual-TTS
with two extra columns:
column
description
audio_filename, text, speaker
unchanged from malaysia-ai/Multilingual-TTS
language
language detected from the text (transcript) column of every row
post-normalized
text after rule-based punctuation / capitalization normalization
All original columns and the file/folder layout are preserved: 1552 subsets /
1561 parquet files, 121,842,371 rows (August 2026). Every… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS-language.genshin-voice-v3.5-mandarin
Dataset Card for Genshin Voice
Dataset Description
Dataset Summary
The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game.
Languages
The text in the dataset is in Mandarin.
Dataset Creation
Source Data
Initial Data Collection and Normalization
The data was obtained by unpacking the Genshin Impact game.
Who are the source language producers?
The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.5-mandarin.audio-v2-transcripts
Overview
This dataset provides full, machine-generated transcriptions for the entire audio-v2 dataset, containing >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license.
It was released on May 18th, 2025.
You can find the full list of sources in this dataset under the audio-v2 dataset's sources.txt.
All files were transcribed using the process.py pipeline, performing:
Frame-level VAD
Machine transcription using ivrit.ai's whisper-large-v3-turbo engine with the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2-transcripts.vivos\
VIVOS is a free Vietnamese speech corpus consisting of 15 hours of recording speech prepared for
Vietnamese Automatic Speech Recognition task.
The corpus was prepared by AILAB, a computer science lab of VNUHCM - University of Science, with Prof. Vu Hai Quan is the head of.
We publish this corpus in hope to attract more scientists to solve Vietnamese speech recognition problems.thai-aligner-bench
Thai Aligner Bench
🚧 Development in progress.
How accurately can a forced aligner place Thai token and word boundaries in
speech? This is a self-contained benchmark: one Python file
(aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing
ground truth. No Thai NLP stack or other code is needed — just
numpy soundfile torch torchaudio transformers.
The ground truth is what makes the dataset useful: the audio was rendered by a
TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.ASR_Code_Switch
ASR Code-Switching Benchmark
A curated benchmark of 1,200 code-switching utterances (300 per language pair)
for evaluating commercial ASR systems on multilingual speech with intra-sentential
language switching.
Paper
Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German
arXiv link
Language pairs
Split
Language pair
Samples
Scripts
egyptian_arabic_english
Egyptian Arabic–English
300
Arabic + Latin… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/ASR_Code_Switch.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.dutch-tts-labeled-complete
Dutch TTS Dataset - Complete Labeled
A comprehensive Dutch text-to-speech dataset with 596,508 audio samples totaling 234GB of audio data.
Quick Preview
The default config shows a 100-row sample for the dataset viewer. To access the full dataset, use the full config.
Dataset Description
This dataset contains Dutch speech recordings with rich metadata including:
Emotion labels (neutral, happy, sad, angry)
Speaker IDs (239,388 unique speakers)… See the full description on the dataset page: https://huggingface.co/datasets/AITRADER/dutch-tts-labeled-complete.STEAK
STEAK — Speech-to-Text for Error of Atc readbacK
STEAK is a synthetically generated dataset of ATCO–pilot radio exchanges
— both the text and the audio are synthetic:
Text: assembled by formal rules, following an ontology of ATCO–pilot
exchanges.
Audio: TTS → voice timbre / accent conversion (seed-vc) →
noise addition (noise captured from real ATCO2 recordings).
One row = one audio (one ATCO controller utterance or one pilot readback).
2,519,694 audios. The ATCO↔pilot pair is… See the full description on the dataset page: https://huggingface.co/datasets/DEEL-AI/STEAK.hailuo-ai-voices
Hailuo AI Voices Dataset 🎤
A curated collection of high-quality voice recordings with corresponding transcriptions and phoneme analysis. This dataset is designed for speech recognition, text-to-speech, and voice analysis tasks.
📊 Dataset Overview
The dataset provides a comprehensive collection of voice samples with the following features:
Feature
Description
Audio Files
High-quality WAV format recordings
Transcription
Accurate transcriptions of each… See the full description on the dataset page: https://huggingface.co/datasets/unlimitedbytes/hailuo-ai-voices.Redmond-Sentence-Recall
Dataset Summary
The Redmond Sentence Recall (RSR) measures a child’s ability to repeat sentences that contain regular past tense forms and past participle forms (e.g., “He kicked” vs. “He was kicked”). This task helps identify language impairments, with each child repeating 16 sentences heard through headphones. The dataset includes anonymized audio recordings of these repetitions.
What makes the RSR dataset uniquely valuable is its focus on sentence recall using both regular past… See the full description on the dataset page: https://huggingface.co/datasets/ai4exceptionaled/Redmond-Sentence-Recall.knesset-plenums-whisper-training
Dataset Card for ivrit.ai - Knesset Plenums Whisper Training
This is a whisper-formatted version of the ivrit.ai Knesset Plenums dataset.
This dataset was created by splitting long audio recordings, along with their respective transcriptions, into audio slices of 30 seconds or less.
Each such slice represents one or more consecutive segments, along with timestamp token data and the previous slice's transcription.
The code for this dataset preparation process is available on the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums-whisper-training.thai-dialect-isan-dataset
Dataset Card for Thai Dialect Isan Speech Corpus
Dataset Description
This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language.
The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai-dialect-isan-dataset.dia-aishell4-test
AISHELL-4 — test split (meeting diarization)
Copie du split test d'AISHELL-4, un corpus de réunions en mandarin capturé
par un array de 8 micros (on garde ici la version single-channel extraite
pour les benchmarks diarisation).
Contenu
20 sessions de réunion (3–7 speakers / session, durée variable)
Audio : FLAC mono
Annotations : RTTM par session
Langue : mandarin (zh)
Licence : Apache-2.0 (upstream)
Structure
dia-aishell4-test/
├── audio/test/<file_id>.flac… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-aishell4-test.Meta_STT_ZH_AIShell3
Meta Speech Recognition Mandarin Dataset (AISHELL3)
This dataset contains both metadata and audio files for Mandarin speech recognition samples from the AISHELL3 corpus.
Dataset Statistics
Splits and Sample Counts
train: 60098 samples
valid: 3163 samples
test: 24772 samples
Example Samples
train
{
"audio_filepath": "/external4/datasets/Mandarin/AISHELL3/wavs_train/SSB00430356.wav",
"text": "她以 ENTITY_PRODUCT 滴鸡精 END 调养身体。 AGE_14_25… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_ZH_AIShell3.a5sv2-asr-benchmark-dataset
A5Sv2 ASR Benchmark Dataset
Public references, saved predictions, scores, and provenance for the
A5Sv2 ASR benchmark. The benchmark evaluates
streaming English ASR on four fixed public corpora with approximately equal normalized reference
word counts.
Corpus
Fixed selection
Reference words
Audio in this repository
Mega-ASR / Voices-in-the-Wild-2M
1,250 utterances, 250 per acoustic condition
32,928
Yes
AMI
7 scenario-only unseen-evaluation meetings
32,928
Yes
DiPCo… See the full description on the dataset page: https://huggingface.co/datasets/AirCaps/a5sv2-asr-benchmark-dataset.parlament_parla_v3
Dataset Card for ParlamentParla v3 - Speech Corpus of Catalan Parliamentary Sessions
A speech corpus composed of Catalan Parliamentary Sessions.The v3 and last version of the corpus includes both clean and other quality segments, divided into short segments (less than 30 seconds) and long segments (more than 30 seconds). The total dataset encompasses 1059h 48m 04s of speech, including 945h 51m 06s for the short segments and 113h 56m 58s for the long segments, with a total of… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/parlament_parla_v3.enhanced-audiosnippets-long-2-8M
Enhanced Audiosnippets Long 2.8M
Enhanced version of mitermix/audiosnippets_long_2_8M with speech enhancement, emotion annotations, speaker embeddings, and comprehensive metadata analysis.
Dataset Summary
Metric
Value
Total samples
2,633,037
Total audio hours
4,932 h
Duration range
3.0s - 1124.3s
Mean duration
6.7s
Audio format
WAV, 48kHz mono
Tar files
1,410
Processing Pipeline
Each audio sample was processed through:
Speech… See the full description on the dataset page: https://huggingface.co/datasets/ai-music4you3/enhanced-audiosnippets-long-2-8M.thai-contextasr-bench
Thai Contextual-Biasing ASR Benchmark
TL;DR
Does your Thai ASR system actually use the context you give it (e.g., a list of names, custom words from your own dictionary) — and does it hallucinate when the context is irrelevant?
Each utterance comes with a bias list: entity strings (brands, person names,
places) that may or may not be spoken in the audio, written the way a real Thai user
would write them — one list, mixed Thai and Latin script. A good system does… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-contextasr-bench.genshin-voice-v3.4-mandarin
Dataset Card for Genshin Voice
Dataset Description
Dataset Summary
The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game.
Languages
The text in the dataset is in Mandarin.
Dataset Creation
Source Data
Initial Data Collection and Normalization
The data was obtained by unpacking the Genshin Impact game.
Who are the source language producers?
The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.4-mandarin.IndicCMix
IndicCMix
Most Indic NLP data assumes people write in one script and one language at a time. Real chat looks nothing like that. You get Hindi words in Roman letters, English verbs in the middle of a Tamil sentence, and the same person switching scripts halfway through a paragraph.
This dataset is an attempt to cover that actual messiness. For every English sentence, you get three different Indic renderings of it: one code-mixed in the native script, one clean native-script… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicCMix.
