datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spanish-slang-stt-data
Spanish Regional Speech-to-Text Dataset
A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models.
Dataset Description
This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions:
Region
Samples
Description
Mexico
17,725
Mexican Spanish including CIEMPIESS corpus
Spain
11,360
Castilian Spanish from TEDx and Common Voice
Argentina
5,839
Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.Seamless_Dummy_Dataset_Fixed_3
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
tr-full-dataset
TR-Full_dataset
This is a merged speech dataset containing 41427 audio segments from 88 source datasets.
Dataset Information
Total Segments: 41427
Speakers: 222
Languages: tr
Emotions: neutral, angry, sad, happy
Original Datasets: 88
Dataset Structure
Each example contains:
audio: Audio file (WAV format, original sampling rate preserved)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr-full-dataset.korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.my-voxtral-datasetContextTTS_dataset
ContextTTS Evaluation Dataset
This is the official evaluation dataset for the paper "[ContextTTS Eval: A Benchmark for Evaluating Long-Form
Contextual Expressive Text-to-Speech]". It is designed to evaluate the performance of multi-modal speech synthesis, specifically focusing on context-aware prosody and timbre consistency in Chinese conversations and audiobooks.
Dataset Summary
The dataset consists of high-quality Chinese audio-text pairs, organized into three distinct… See the full description on the dataset page: https://huggingface.co/datasets/rodenhhh/ContextTTS_dataset.MCF-Dataset
MCF: Text LLMS For Multimodal Emotional Causality
Data
Dataset task definition and annotation example of the MCF framework. The framework contains two core subtasks:
five-tuple
element extraction (identifying Target, Holder, Aspect, Opinion, Sentiment, and Rationale) and sentiment chain
analysis (constructing causal
relationship chains between emotional events).
The dataset is provided with the following structure. Each sample includes video, audio, and dialogue… See the full description on the dataset page: https://huggingface.co/datasets/ZHANGYUXUAN-zR/MCF-Dataset.personaplex-finetuning-pharma-data-sample
PersonaPlex Finetuning — Pharma Data Sample
A 10-example slice of the synthetic patient-support / medication
adherence dataset used to train
demegire/personaplex-finetune-pharma.
The on-disk layout below is exactly what the trainer in
emotion-machine-org/personaplex-finetune
consumes — use this as a template when building your own.
Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at
sample scale).
Layout
.
├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.naijavoices_dataset_85_hours_tts_bestFull dataset re-upload with more statistics as well as filtering scripts that give the top x files or best x hours for tts based and calculations from acoustic metrics\color{Blue}{\large \textbf{Full dataset re-upload with more statistics as well as filtering scripts that give the top x files or best x hours for tts based and calculations from acoustic… See the full description on the dataset page: https://huggingface.co/datasets/David-A-Amoo/naijavoices_dataset_85_hours_tts_best.song_dataset
🎵 Vietnamese Song Lyrics and Word Timestamps Dataset
Dataset Summary
The song_dataset provides high-quality Vietnamese song data, including metadata, full lyrics, and particularly word-level timestamps.
This dataset is optimally designed for tasks such as:
Training and evaluating automatic speech recognition (ASR) models on music.
Lyrics synchronization (Lyrics Alignment / Karaoke generation).
Natural language processing (NLP) analysis on song lyrics.
The data is… See the full description on the dataset page: https://huggingface.co/datasets/sunbv56/song_dataset.jeju_potato_datasetsmascarade-dsp-dataset
Mascarade — DSP & Signal Processing Q&A
✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11)
Per-sample Stack Exchange Electronics attribution recovered via the SE
/search/advanced + /questions/{id} API search :
169 samples (~5.35 %) confirmed as Stack Exchange Electronics
(CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution
(URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60).
535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/electron-rare/mascarade-dsp-dataset.ns-urdu-datasetbhashini-datasetchlid-datasetsmall-german-medical-dialogue-dataset-for-moshi
Small german dialogue dataset
This dataset contains 500 completely made up medical phonecall dialogues between patients and a GP's office.
Dataset Details
Dataset Description
500 made up phonecalls that were first created with AI as text.
The audio was then created using Openai tts-1-hd and the accurately timestamped transcripts were added.
The audio files are formatted like this:
Stereo with split channels:
Speaker A is on the left channel… See the full description on the dataset page: https://huggingface.co/datasets/chtugha/small-german-medical-dialogue-dataset-for-moshi.stage1a_smoke_data
stage1a_smoke_data — AuT-ready 128-mel TFRecords (en/zh)
Smoke-scale training data for Stage 1A input audio alignment of a Qwen3-ASR-AuT → MLP →
frozen-VL-LLM omni model. Audio is pre-extracted 128-bin log-mel (the Qwen3-ASR AuT frontend:
WhisperFeatureExtractor, 16 kHz, hop 160, n_fft 400) so training only needs to run the frozen AuT
encoder — no raw-audio decoding at train time.
113,396 samples across 4 sources, stored as GZIP-compressed TFRecords (one file per source shard).… See the full description on the dataset page: https://huggingface.co/datasets/Letian2003/stage1a_smoke_data.synthetic-patient-dr-data
Synthetic Patient DR Data
Synthetic doctor-patient consultation dataset with structured clinical outputs and optional full-consultation audio.
Dataset Summary
This dataset was generated for research and prototyping in:
clinical dialogue generation
structured clinical extraction
text-to-audio workflows
conversational healthcare modeling
All consultations are synthetic and should not be treated as real clinical encounters.
Export Metadata
Mode: audio
Repo… See the full description on the dataset page: https://huggingface.co/datasets/TumeloKonaite/synthetic-patient-dr-data.song_dataset_chunked
Vietnamese Songs Word-Level Timestamp Dataset (Chunked)
This dataset contains word-level timestamp information for Vietnamese songs, specifically pre-chunked into segments up to 30 seconds for use in training or fine-tuning speech recognition (ASR) systems like Whisper.
Dataset Summary
The song_dataset_chunked provides high-quality Vietnamese song data, properly segmented into optimal ~30-second sequences.
Duration Insights:
Train split (train_chunked.jsonl): ~ 230.62… See the full description on the dataset page: https://huggingface.co/datasets/sunbv56/song_dataset_chunked.Onomatopoeia_Dataset🎧 Onomatopoeia Dataset (Audio → Manga Expression)
音声解析結果をもとに、日本語のオノマトペ(擬音語・擬態語)を生成するためのデータセットです。
本データセットは、音そのものではなく、音から推定された特徴・空間・情景を入力とする構造化データであり、
漫画的な表現生成を目的としたマルチモーダルデータです。
📌 Dataset Summary
本データセットは以下のパイプラインから生成されています:
Audio
↓
Audio Features (04_features.json)
↓
Audio Events (05_audio_events.json)
↓
Space Judgement (06_space_judgement.json)
↓
Scene Interpretation (07_scene_interpretation.json)
↓
Onomatopoeia (08_onomatopoeia.json)
👉 音 → 空間 → 情景 → オノマトペ
という段階的生成構造を持ちます。
📊… See the full description on the dataset page: https://huggingface.co/datasets/yadorigi/Onomatopoeia_Dataset.concept-datasetSPS-Bopha-Voice-Dataset-v1
VibeVoice Fine-Tuning Dataset: SPS-Bopha-Voice-Dataset-v1
This dataset is formatted for fine-tuning VibeVoice.
Structure
training_data.jsonl: The main manifest file containing transcriptions and paths.
chunks_staging/: Directory containing the audio clips.
Usage with VibeVoice
Clone this repository:
git clone https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1
cd SPS-Bopha-Voice-Dataset-v1
Run the training script pointing to… See the full description on the dataset page: https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1.scasr_datasetmascarade-dsp-dataset
Mascarade — DSP & Signal Processing Q&A
✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11)
Per-sample Stack Exchange Electronics attribution recovered via the SE
/search/advanced + /questions/{id} API search :
169 samples (~5.35 %) confirmed as Stack Exchange Electronics
(CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution
(URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60).
535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/mascarade-dsp-dataset.fma-dataset-aug-caption
FMA-CLAP Caption Augmentation Dataset
Overview
This dataset is an enhanced version of the FMA (Free Music Archive) dataset, where we have augmented the original metadata with natural language captions generated using the CLAP (Contrastive Language-Audio Pretraining) model. The captions describe the genre, style, mood, and instrumentation of each track, making it more suitable for zero-shot learning, music classification, and text-to-music generation tasks.… See the full description on the dataset page: https://huggingface.co/datasets/solbon1212/fma-dataset-aug-caption.Indic_New_dataset_TTS
Indic TTS Dataset Hub (Mozilla)
Validated audio–text pairs for multiple Indic languages from Mozilla Common Voice.
Select the language from the Subset dropdown in the Dataset Viewer.
Columns
audio: WAV audio clip (16kHz)
text: transcription
duration: length in seconds
speaking_rate: characters per second
sopho-poetry-tts-data
sopho-poetry-tts-data
Training data for the SuFei on-device Chinese-poetry TTS pipeline
(sopho-poetry-tts-train).
Two co-equal lineages, one per teacher model — both active (the FS2 set is
the provenance of production v9 and stays reusable).
Lineage
Teacher
Trained
Poems
cosyvoice3/
Fun-CosyVoice3-0.5B (Apache-2.0)
m3_v6
1023
paddlespeech_fs2/
PaddleSpeech FS2 CSMSC (Apache-2.0)
v9
320
Layout
poems.jsonl # shared source text, keyed… See the full description on the dataset page: https://huggingface.co/datasets/Sopho/sopho-poetry-tts-data.topipl-spanish-datasetvoice-dataset
Voice Dataset
Collected from the web uploader tool.
Voice Dataset
Collected from the web uploader tool.
dataflow-mm-audio_asr_pipeline
