datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
StreamAudio-2M
StreamAudio-2M
Large-scale streaming-audio dataset for audio-LLM / audio-agent training. Each row is a
stream: a sequence of audio turns sharing one unified schema. ~2.28M unique audio clips
are organised into six task subsets.
Subsets
Subset
Rows
Description
Stream_Audio_Understanding
90,738
Montages of audio-understanding clips (AudioSet / FMA): captions, choice & open QA
Real_time_ASR
28,109
Streams of ASR clips (CommonVoice / GigaSpeech /… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/StreamAudio-2M.spanish-slang-stt-data
Spanish Regional Speech-to-Text Dataset
A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models.
Dataset Description
This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions:
Region
Samples
Description
Mexico
17,725
Mexican Spanish including CIEMPIESS corpus
Spain
11,360
Castilian Spanish from TEDx and Common Voice
Argentina
5,839
Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.verbalyze-stt-bench
Verbalyze: Indic Speech & ITN Benchmark (12 Languages)
Verbalyze is a scenario-weighted benchmark dataset designed to evaluate and train Speech-to-Text (ASR) and Inverse Text Normalization (ITN) models on real-world Indic speech phenomena.
It covers 12 major Indian languages with 172,800 balanced utterances categorized across 11 edge-case scenarios where standard speech models typically fail.
Languages Covered
Language
Code
Samples
Script
Assamese
as
14… See the full description on the dataset page: https://huggingface.co/datasets/ansh-rohilla/verbalyze-stt-bench.star-wars-dataset
Star Wars dialogue dataset - annotation layers (snapshot 2026-09-21)
One master cue table serving diarization, ASR and dialogue-LLM training: every subtitle cue part of
15 titles with a forced-aligned time span, a character label and provenance. No audio or video is
included - the source media is copyrighted. Clip paths in the manifests resolve once you rebuild the clips
from your own copies with the pipeline code (export_asr.py, export_diarization.py).
The text layers contain… See the full description on the dataset page: https://huggingface.co/datasets/IMJONEZZ/star-wars-dataset.pseudolabel-malaya-speech-stt-train-whisper-large-v3advanced-soundscapes-stage-1
Advanced Soundscapes Stage 1 — Raw Components
This dataset contains Stage 1 output from the LAION Universal Audio Annotation Pipeline (UAAP) data generation plan.
Contents
0 shard(s) containing 0 soundscape recipes with raw audio components
Each soundscape row includes:
recipe.json — full recipe with timeline, events, loudness, speaker IDs, overlap/density settings
spkN.flac / spkN.json — raw speech components + full source metadata
musicN.flac / musicN.json —… See the full description on the dataset page: https://huggingface.co/datasets/ChristophSchuhmann/advanced-soundscapes-stage-1.stage1a_smoke_data
stage1a_smoke_data — AuT-ready 128-mel TFRecords (en/zh)
Smoke-scale training data for Stage 1A input audio alignment of a Qwen3-ASR-AuT → MLP →
frozen-VL-LLM omni model. Audio is pre-extracted 128-bin log-mel (the Qwen3-ASR AuT frontend:
WhisperFeatureExtractor, 16 kHz, hop 160, n_fft 400) so training only needs to run the frozen AuT
encoder — no raw-audio decoding at train time.
113,396 samples across 4 sources, stored as GZIP-compressed TFRecords (one file per source shard).… See the full description on the dataset page: https://huggingface.co/datasets/Letian2003/stage1a_smoke_data.hoiku-yougo-stt-ja
保育用語STT最適化データセット
Japanese Childcare Vocabulary Dataset for STT Optimization
🇯🇵 日本語
概要
このデータセットは、日本の保育現場における音声認識(STT)の精度向上を目的として構築された、保育専用の構造化語彙集です。
一般的なSTTエンジン(Google・Amazon・OpenAI Whisper等)は、保育現場で日常的に使われる専門用語を正確に認識できないという問題があります。たとえば:
実際の発話
STTが誤認識する例
誤飲(ごいん)
語韻・誤音
午睡(ごすい)
誤推・誤睡
降園(こうえん)
公演・後援
点呼(てんこ)
転校・天候
視診(ししん)
試診・指診
このデータセットは、こうした保育特有の同音異義語・誤変換パターンを体系的に収録し、STTエンジンのカスタム辞書・RAG(検索拡張生成)・AIコール(音声電話応答)に活用できるよう設計されています。… See the full description on the dataset page: https://huggingface.co/datasets/IDEMITSU/hoiku-yougo-stt-ja.
