datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Seamless_Dummy_Dataset_Fixed_3
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
tr-full-dataset
TR-Full_dataset
This is a merged speech dataset containing 41427 audio segments from 88 source datasets.
Dataset Information
Total Segments: 41427
Speakers: 222
Languages: tr
Emotions: neutral, angry, sad, happy
Original Datasets: 88
Dataset Structure
Each example contains:
audio: Audio file (WAV format, original sampling rate preserved)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr-full-dataset.korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.my-voxtral-datasetContextTTS_dataset
ContextTTS Evaluation Dataset
This is the official evaluation dataset for the paper "[ContextTTS Eval: A Benchmark for Evaluating Long-Form
Contextual Expressive Text-to-Speech]". It is designed to evaluate the performance of multi-modal speech synthesis, specifically focusing on context-aware prosody and timbre consistency in Chinese conversations and audiobooks.
Dataset Summary
The dataset consists of high-quality Chinese audio-text pairs, organized into three distinct… See the full description on the dataset page: https://huggingface.co/datasets/rodenhhh/ContextTTS_dataset.MCF-Dataset
MCF: Text LLMS For Multimodal Emotional Causality
Data
Dataset task definition and annotation example of the MCF framework. The framework contains two core subtasks:
five-tuple
element extraction (identifying Target, Holder, Aspect, Opinion, Sentiment, and Rationale) and sentiment chain
analysis (constructing causal
relationship chains between emotional events).
The dataset is provided with the following structure. Each sample includes video, audio, and dialogue… See the full description on the dataset page: https://huggingface.co/datasets/ZHANGYUXUAN-zR/MCF-Dataset.naijavoices_dataset_85_hours_tts_bestFull dataset re-upload with more statistics as well as filtering scripts that give the top x files or best x hours for tts based and calculations from acoustic metrics\color{Blue}{\large \textbf{Full dataset re-upload with more statistics as well as filtering scripts that give the top x files or best x hours for tts based and calculations from acoustic… See the full description on the dataset page: https://huggingface.co/datasets/David-A-Amoo/naijavoices_dataset_85_hours_tts_best.song_dataset
🎵 Vietnamese Song Lyrics and Word Timestamps Dataset
Dataset Summary
The song_dataset provides high-quality Vietnamese song data, including metadata, full lyrics, and particularly word-level timestamps.
This dataset is optimally designed for tasks such as:
Training and evaluating automatic speech recognition (ASR) models on music.
Lyrics synchronization (Lyrics Alignment / Karaoke generation).
Natural language processing (NLP) analysis on song lyrics.
The data is… See the full description on the dataset page: https://huggingface.co/datasets/sunbv56/song_dataset.jeju_potato_datasetsmascarade-dsp-dataset
Mascarade — DSP & Signal Processing Q&A
✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11)
Per-sample Stack Exchange Electronics attribution recovered via the SE
/search/advanced + /questions/{id} API search :
169 samples (~5.35 %) confirmed as Stack Exchange Electronics
(CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution
(URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60).
535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/electron-rare/mascarade-dsp-dataset.ns-urdu-datasetbhashini-datasetchlid-datasetsmall-german-medical-dialogue-dataset-for-moshi
Small german dialogue dataset
This dataset contains 500 completely made up medical phonecall dialogues between patients and a GP's office.
Dataset Details
Dataset Description
500 made up phonecalls that were first created with AI as text.
The audio was then created using Openai tts-1-hd and the accurately timestamped transcripts were added.
The audio files are formatted like this:
Stereo with split channels:
Speaker A is on the left channel… See the full description on the dataset page: https://huggingface.co/datasets/chtugha/small-german-medical-dialogue-dataset-for-moshi.song_dataset_chunked
Vietnamese Songs Word-Level Timestamp Dataset (Chunked)
This dataset contains word-level timestamp information for Vietnamese songs, specifically pre-chunked into segments up to 30 seconds for use in training or fine-tuning speech recognition (ASR) systems like Whisper.
Dataset Summary
The song_dataset_chunked provides high-quality Vietnamese song data, properly segmented into optimal ~30-second sequences.
Duration Insights:
Train split (train_chunked.jsonl): ~ 230.62… See the full description on the dataset page: https://huggingface.co/datasets/sunbv56/song_dataset_chunked.Onomatopoeia_Dataset🎧 Onomatopoeia Dataset (Audio → Manga Expression)
音声解析結果をもとに、日本語のオノマトペ(擬音語・擬態語)を生成するためのデータセットです。
本データセットは、音そのものではなく、音から推定された特徴・空間・情景を入力とする構造化データであり、
漫画的な表現生成を目的としたマルチモーダルデータです。
📌 Dataset Summary
本データセットは以下のパイプラインから生成されています:
Audio
↓
Audio Features (04_features.json)
↓
Audio Events (05_audio_events.json)
↓
Space Judgement (06_space_judgement.json)
↓
Scene Interpretation (07_scene_interpretation.json)
↓
Onomatopoeia (08_onomatopoeia.json)
👉 音 → 空間 → 情景 → オノマトペ
という段階的生成構造を持ちます。
📊… See the full description on the dataset page: https://huggingface.co/datasets/yadorigi/Onomatopoeia_Dataset.concept-datasetSPS-Bopha-Voice-Dataset-v1
VibeVoice Fine-Tuning Dataset: SPS-Bopha-Voice-Dataset-v1
This dataset is formatted for fine-tuning VibeVoice.
Structure
training_data.jsonl: The main manifest file containing transcriptions and paths.
chunks_staging/: Directory containing the audio clips.
Usage with VibeVoice
Clone this repository:
git clone https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1
cd SPS-Bopha-Voice-Dataset-v1
Run the training script pointing to… See the full description on the dataset page: https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1.scasr_datasetmascarade-dsp-dataset
Mascarade — DSP & Signal Processing Q&A
✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11)
Per-sample Stack Exchange Electronics attribution recovered via the SE
/search/advanced + /questions/{id} API search :
169 samples (~5.35 %) confirmed as Stack Exchange Electronics
(CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution
(URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60).
535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/mascarade-dsp-dataset.fma-dataset-aug-caption
FMA-CLAP Caption Augmentation Dataset
Overview
This dataset is an enhanced version of the FMA (Free Music Archive) dataset, where we have augmented the original metadata with natural language captions generated using the CLAP (Contrastive Language-Audio Pretraining) model. The captions describe the genre, style, mood, and instrumentation of each track, making it more suitable for zero-shot learning, music classification, and text-to-music generation tasks.… See the full description on the dataset page: https://huggingface.co/datasets/solbon1212/fma-dataset-aug-caption.Indic_New_dataset_TTS
Indic TTS Dataset Hub (Mozilla)
Validated audio–text pairs for multiple Indic languages from Mozilla Common Voice.
Select the language from the Subset dropdown in the Dataset Viewer.
Columns
audio: WAV audio clip (16kHz)
text: transcription
duration: length in seconds
speaking_rate: characters per second
topipl-spanish-datasetvoice-dataset
Voice Dataset
Collected from the web uploader tool.
Voice Dataset
Collected from the web uploader tool.
khmer_speech_news_dataset
Khmer Speech Dataset Processing
This repository contains scripts and instructions for preparing a Khmer speech dataset for machine learning tasks, such as automatic speech recognition (ASR). It demonstrates how to process a collection of audio files and metadata, and save them as Parquet files for efficient use in your training pipelines—without needing torchcodec.
All dataset audio and transcripts in this project are sourced from https://wmc.org.kh/, the official website of… See the full description on the dataset page: https://huggingface.co/datasets/vichetkao/khmer_speech_news_dataset.chester_bennington_30tracks_datasetaitf-dfk3-synthetic-audio-datasetindicvc-datasets2o-dataset
