datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
StreamAudio-2M
StreamAudio-2M
Large-scale streaming-audio dataset for audio-LLM / audio-agent training. Each row is a
stream: a sequence of audio turns sharing one unified schema. ~2.28M unique audio clips
are organised into six task subsets.
Subsets
Subset
Rows
Description
Stream_Audio_Understanding
90,738
Montages of audio-understanding clips (AudioSet / FMA): captions, choice & open QA
Real_time_ASR
28,109
Streams of ASR clips (CommonVoice / GigaSpeech /… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/StreamAudio-2M.starrail-voice
StarRail Voice
StarRail Voice is a dataset of voice lines from the popular game Honkai: Star Rail.
Hugging Face 🤗 StarRail-Voice
ModelScope StarRail-Voice
Last update at 2026-07-16, game version 4.4.0
403437 wavs
60164 without speaker (15%)
61375 without transcription (15%)
57869 without inGameFilename (14%)
Dataset Details
Dataset Description
The dataset contains voice lines from the game's characters in multiple languages, including Chinese… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/starrail-voice.Malaysian-STT-Whisper
Malaysian STT Whisper format
Heavy postprocessing and post-translation to improve pseudolabeled Whisper Large V3. Also include word level timestamp.
Postprocessing
Check repetitive trigrams.
Verify Voice Activity using Silero-VAD.
Verify scores using Force Alignment.
Post-translation
We use mesolitica/nanot5-base-malaysian-translation-v2.1.
Dataset involved
Malaysian context v2
Singaporean context
Indonesian context
Mandarin audio
Tamil audio… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper.arabic-audio-collection-algerian-loubna-stories
Loubna Stories Arabic Speech Dataset
Dataset Summary
The Loubna Stories Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 237 hours of speech recordings and corresponding transcripts.
What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-loubna-stories.Agri_STT_Benchmarking_Dataset
Agri STT Benchmarking Dataset
10,808 farmer voice queries in Hindi, Telugu and Odia, with reference transcripts, for benchmarking automatic speech recognition in agricultural contexts. The audio is included in this repository.
Every recording is a smallholder farmer speaking a question to Farmer.Chat, an AI advisory service run by Digital Green. Reference transcripts were produced by human annotators. Nothing here is read from a script or recorded in a studio, so the audio… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/Agri_STT_Benchmarking_Dataset.sTinyStories
sTinyStories
A spoken version of TinyStories Synthesized with LJ voice using FastSpeech2.
The dataset was synthesized to boost the training of Speech Language Models as detailed in the paper "Slamming: Training a Speech Language Model on One GPU in a Day".
It was first suggested by Cuervo et. al 2024.
We refer you to the SlamKit codebase to see how you can train a SpeechLM with this dataset.
Usage
from datasets importload_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/slprl/sTinyStories.Uzbek-STT-Dataset-780h
Uzbek STT Dataset (~780 hours)
A large Uzbek speech-to-text dataset for training and fine-tuning automatic
speech recognition (ASR) models such as Whisper.
Dataset summary
Language
Uzbek (uz)
Examples
122,464
Total audio
~780 hours
Clip length
up to 30 seconds each
Columns
audio, transcription
Audio
embedded in parquet, original sample rates (e.g. 44.1 kHz / 16 kHz), mono/stereo as recorded
Split
single train split (split it yourself as… See the full description on the dataset page: https://huggingface.co/datasets/Abduqayum/Uzbek-STT-Dataset-780h.talkbank_4_stt
Dataset Card
Dataset Description
This dataset is a benchmark based on the TalkBank[1] corpus—a large multilingual repository of conversational speech that captures real-world, unstructured interactions. We use CA-Bank [2], which focuses on phone conversations between adults, which include natural speech phenomena such as laughter, pauses, and interjections. To ensure the dataset is highly accurate and suitable for benchmarking conversational ASR systems, we employ… See the full description on the dataset page: https://huggingface.co/datasets/diabolocom/talkbank_4_stt.spanish-slang-stt-data
Spanish Regional Speech-to-Text Dataset
A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models.
Dataset Description
This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions:
Region
Samples
Description
Mexico
17,725
Mexican Spanish including CIEMPIESS corpus
Spain
11,360
Castilian Spanish from TEDx and Common Voice
Argentina
5,839
Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.STEAK
STEAK — Speech-to-Text for Error of Atc readbacK
STEAK is a synthetically generated dataset of ATCO–pilot radio exchanges
— both the text and the audio are synthetic:
Text: assembled by formal rules, following an ontology of ATCO–pilot
exchanges.
Audio: TTS → voice timbre / accent conversion (seed-vc) →
noise addition (noise captured from real ATCO2 recordings).
One row = one audio (one ATCO controller utterance or one pilot readback).
2,519,694 audios. The ATCO↔pilot pair is… See the full description on the dataset page: https://huggingface.co/datasets/DEEL-AI/STEAK.Kurisu_Voice
Makise Kurisu Multilingual Voice Dataset
13,999 labelled clips (17.1423 hours) of Makise Kurisu,
including the Amadeus Kurisu variant, in six languages, cut from the
STEINS;GATE games, the anime, and three character songs.
Every clip carries the transcript, a measured acoustic profile, an independent
speaker-identity check, and — where it could be earned rather than guessed — an
expressive tag. This is an unofficial, fan-made dataset with no affiliation to
the STEINS;GATE rights… See the full description on the dataset page: https://huggingface.co/datasets/starrydark/Kurisu_Voice.Meta_STT_ZH_AIShell3
Meta Speech Recognition Mandarin Dataset (AISHELL3)
This dataset contains both metadata and audio files for Mandarin speech recognition samples from the AISHELL3 corpus.
Dataset Statistics
Splits and Sample Counts
train: 60098 samples
valid: 3163 samples
test: 24772 samples
Example Samples
train
{
"audio_filepath": "/external4/datasets/Mandarin/AISHELL3/wavs_train/SSB00430356.wav",
"text": "她以 ENTITY_PRODUCT 滴鸡精 END 调养身体。 AGE_14_25… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_ZH_AIShell3.Meta_STT_HI_Set1
Meta Speech Recognition Hindi Dataset (Set 1)
This dataset contains both metadata and audio files for Hindi speech recognition samples, curated from multiple sources.
Dataset Sources and Credits
This dataset combines samples from the following sources:
AI4Bharat Indic Speech Dataset
Source: https://ai4bharat.org/indic-speech-dataset
License: CC-BY 4.0
Citation: Please cite the original paper if you use this data
Common Voice Hindi
Source:… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_HI_Set1.riigikogu-audio-stenograms-2018-2025
Riigikogu Stenograms 2018-2025
This dataset contains stenograms (approximate transcripts) of the sessions of the Estonian Parliament Riigikogu, spanning a time period from the very end of 2017 to May 2025, together with the corresponding audio.
The transcripts are not verbatim (word-by-word) transcripts but are edited for readability and grammatical correctness. Sentence start and end times (w.r.t. to the corresponding audio file) are provided.
The transcripts are converted into… See the full description on the dataset page: https://huggingface.co/datasets/TalTechNLP/riigikogu-audio-stenograms-2018-2025.advanced-soundscapes-stage-1
Advanced Soundscapes Stage 1 — Raw Components (5M)
This dataset contains Stage 1 output from the LAION Universal Audio Annotation Pipeline (UAAP) data generation plan.
Contents
5,000 shards containing 5,000,000 soundscape recipes with raw audio components
Each soundscape row includes:
recipe.json — full recipe with timeline, events, loudness, speaker IDs, overlap/density settings
spkN.flac / spkN.json — raw speech components + full source metadata
musicN.flac /… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/advanced-soundscapes-stage-1.MultiMed-ST
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
📘 EMNLP 2025
Khai Le-Duc*, Tuyen Tran*, Bach Phan Tat, Nguyen Kim Hai Bui, Quan Dang, Hung-Phong Tran, Thanh-Thuy Nguyen, Ly Nguyen, Tuan-Minh Phan, Thi Thu Phuong Tran, Chris Ngo, Nguyen X. Khanh**, Thanh Nguyen-Tang**
*Equal contribution | **Equal supervision
⭐ If you find this work useful, please consider starring the repo and… See the full description on the dataset page: https://huggingface.co/datasets/leduckhai/MultiMed-ST.ePark_tu_hua_gu_shi_pian_picture_story
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_tu_hua_gu_shi_pian_picture_story
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_tu_hua_gu_shi_pian_picture_story.shenava-koochik-number-stress-36k
Shenava Koochik Number Stress 36K
This repository contains the 36,000-row Persian number-stress text curriculum and
an audio-backed subset of 6,669 short clips generated with Gemini TTS.
The Dataset Viewer default configuration is the 6,669-clip audio subset. It
exposes audio, text, category, and duration_s; the metadata links each
row to its WAV using a relative file_name such as audio/num-000000.wav.
The complete text-only source remains available at
data/train.jsonl (36,000… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/shenava-koochik-number-stress-36k.Meta_STT_EN_Set2
Meta Speech Recognition English Dataset (Set 2)
This dataset contains both metadata and audio files for English speech recognition samples.
Dataset Statistics
Splits and Sample Counts
train: 42961 samples
valid: 2387 samples
test: 2387 samples
Example Samples
train
{
"audio_filepath": "/external1/datasets/asr-himanshu/avspeech-data/audio/AzSutepklXI_2.wav",
"text": "To Jesus, so God is faithful, because when he keeps, you know, when… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/Meta_STT_EN_Set2.stt-unified-bench
am-pranav/stt-unified-bench
Private, language/locale-partitioned mini-benchmark for STT models.
Each subset is a dataset config (e.g., en, de, en_indian_accent, hi_in) with a single split val.
Audio is staged at 16 kHz and stored in-repo for reproducibility.
Schema
audio : Audio(sampling_rate=16000, decode=False)
text : reference transcription
lang : implied by dataset config name
source : upstream dataset tag
id : source-stable id
⚠️ For internal evaluation only.… See the full description on the dataset page: https://huggingface.co/datasets/am-pranav/stt-unified-bench.MCIF-ST
MCIF-ST: Context-aware Speech Recognition and Speech Translation from MCIF
MCIF-ST provides both long-form and short-form ready-to-use
Automatic Speech Recogniton (ASR) and Speech Translation (ST) data derived
from MCIF (Multimodal
Crosslingual Instruction Following), a multilingual benchmark based on
scientific talks. While the original MCIF release packages its content as
instruction-following rows (multimodal context + prompt + expected
answer, for… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MCIF-ST.khmer-speech-dataset
Khmer ASR Cultural Dataset
727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided.
Language: Khmer (khm).
Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata.
Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page: https://huggingface.co/datasets/Sophy15-St/khmer-speech-dataset.Zeroth-STT-Korean
Zeroth-STT-Korean Dataset
Description
This is a shuffled version of the Zeroth-STT-Ko dataset.
Citation
Zeroth-Korean Dataset, created by [Lucas Jo(@Atlas Guide Inc.) and Wonkyum Lee(@Gridspace Inc.)], 2023.
Available at https://github.com/goodatlas/zeroth under CC-BY-4.0 license.
Junhoee/STT_Korean_Dataset_80000 Dataset, created by [Junhoee], 2024.
Available at https://huggingface.co/datasets/Junhoee/STT_Korean_Dataset_80000
STOMA
STOMA: A Multi-Speaker Greek Speech Corpus
STOMA is a new multi-speaker Greek speech corpus designed to advance research in text-to-speech (TTS) synthesis and related speech technologies for Greek, an under-resourced language. The corpus comprises approximately 23 hours of studio-recorded read speech from six native speakers (three male and three female), captured under controlled studio conditions using a dual-booth setup to ensure acoustic consistency and high signal quality. The… See the full description on the dataset page: https://huggingface.co/datasets/aangelakis/STOMA.honkai-star-rail-voices
Honkai: Star Rail — Voice Lines (Multi-Language)
An archive of character voice data extracted from Honkai: Star Rail (崩壊:スターレイル), repackaged as Parquet shards per audio language.
Dataset Summary
Field
Value
Game
Honkai: Star Rail (崩壊:スターレイル)
Publisher
HoYoverse / miHoYo Co., Ltd.
Languages
中文 (zh), 日本語 (ja), English (en), 한국어 (ko)
Game version
3.8
Source format
WAV + sidecar transcripts (.lab / .txt)
Distribution format
Apache Parquet (zstd), ~500 MiB… See the full description on the dataset page: https://huggingface.co/datasets/ultemica/honkai-star-rail-voices.zeroth-STT-Ko
Zeroth-STT-Ko Dataset
Description
This dataset combines the following publicly available Korean language datasets:
Junhoee/STT_Korean_Dataset_80000
and
Zeroth-Korean Dataset (from Project: Zeroth, by GoodAtlas and Gridspace)
This provides over 102K rows of data (sentences) in total.
Citation
Zeroth-Korean Dataset, created by [Lucas Jo(@Atlas Guide Inc.) and Wonkyum Lee(@Gridspace Inc.)], 2023.
Available at https://github.com/goodatlas/zeroth under CC-BY-4.0… See the full description on the dataset page: https://huggingface.co/datasets/o0dimplz0o/zeroth-STT-Ko.stt-sampler-v1
stt-sampler-v1
Licensing: clips inherit their source dataset's license — CC-BY-4.0
for MInDS-14 and FLEURS clips, CC-BY-NC-SA-4.0 for Speech-MASSIVE
clips (source_dataset column identifies each clip's origin).
A small, balanced, representative multilingual ASR eval sampler for the
OVOS Plugin Arena:
100 clips per language x 20 locales = 2000 clips, 16 kHz mono float32,
one config per locale (load_dataset("OpenVoiceOS/stt-sampler-v1", "<lang>")).
Designed to seed every STT… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/stt-sampler-v1.steinsgate-voice
STEINS;GATE Voice Dataset
A voice dataset with transcriptions extracted from the STEINS;GATE visual novel series.
Information
This dataset contains official annotations from the game, including speaker name and transcription.
Current Status:
✅ STEINS;GATE
⏳ STEINS;GATE 0
⏳ STEINS;GATE: My Darling's Embrace
⏳ STEINS;GATE: Linear Bounded Phenogram
Uses
from datasets import load_dataset
# This will stream all data from all games
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/ntaquan0125/steinsgate-voice.DASS2019_NLP
Dataset Card for DASS2019_NLP
This dataset contains audio and transcript content from DASS2019, the manually transcribed version of the Digital Archive of Southern Speech.
It may be suitable for speech-related NLP processing, modelling, and fine-tuning tasks.
Dataset Details
DASS (Kretzschmar et al. 2012) comprises dialectological interviews with 64 informants conducted between 1968 and 1983; it is a subset of the larger Linguistic Atlas of the Gulf States (LAGS… See the full description on the dataset page: https://huggingface.co/datasets/stcoats/DASS2019_NLP.mongolian-stt-dataset
Mongolian Speech Dataset (v24 corpus)
Mongolian (Cyrillic Khalkha) read speech for ASR fine-tuning: 146.9 hours
across Common Voice v24, FLEURS, and MBSpeech.
2026-07-30 — two changes, read this if you pulled before that date.
YouTube-sourced audio removed. 598 clips (559 train / 39 validation, ~1.1 h)
are gone. Every remaining row is read speech from a redistributable public corpus.
This repo now hosts the v24 corpus. It previously held the v20 blend
(57,320 train / 3,017… See the full description on the dataset page: https://huggingface.co/datasets/Blgn94/mongolian-stt-dataset.
