datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GTSinger
GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
Yu Zhang*, Changhao Pan*, Wenxiang Guo*, Ruiqi Li, Zhiyuan Zhu, Jialei Wang, Wenhao Xu, Jingyu Lu, Zhiqing Hong, Chuxin Wang, LiChao Zhang, Jinzheng He, Ziyue Jiang, Yuxin Chen, Chen Yang, Jiecheng Zhou, Xinyu Cheng, Zhou Zhao | Zhejiang University
Dataset of GTSinger (NeurIPS 2024 Spotlight): A Global Multi-Technique Singing Corpus with Realistic Music Scores for All… See the full description on the dataset page: https://huggingface.co/datasets/AaronZ345/GTSinger.vibevoice-gptinformal_persian-single-speakerAudioMCQ-StrongAC-GeminiCoT
[ICLR 2026] [DCASE 2026 Training Set] AudioMCQ-StrongAC-GeminiCoT
This dataset is a highly curated subset of the AudioMCQ dataset, containing samples with native Chain-of-Thought (CoT) reasoning from Gemini 3.1 Pro that were answered correctly.
Note: The native CoT reasoning has been summarized by Gemini's internal algorithm before output, yet still retains rich audio details including timestamps, acoustic descriptions, and step-by-step temporal analysis.
🏆 DCASE 2026… See the full description on the dataset page: https://huggingface.co/datasets/Harland/AudioMCQ-StrongAC-GeminiCoT.GSgdpval_preference_rubricsAVE-Speech
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
Abstract
AVE Speech is a large-scale Mandarin speech corpus that pairs synchronized audio, lip video and surface electromyography (EMG) recordings. The dataset contains 100 sentences read by 100 native speakers. Each participant repeated the full corpus ten times, yielding over 55 hours of data per modality. These complementary signals enable… See the full description on the dataset page: https://huggingface.co/datasets/MML-Group/AVE-Speech.multichannel-meetings-10h
GroundTruth Multi-Channel Meeting Audio Dataset (10h)
Summary
This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant.
Each meeting includes:
One full meeting recording (room microphone)
Individual close-talk recordings for each participant (one file per speaker)
Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.chinese-lips-speech-slide-probe
Chinese-LiPS Speech + Slide Probe
A self-contained probe set for testing whether visual slide context helps
simultaneous speech translation — with the input as audio, not transcripts.
Why audio matters: feeding a transcript to a text LLM deletes the acoustic
ambiguity (homophones, polysemy) that slide context is meant to resolve; the
transcript already commits to one reading. Any honest test of "does vision help
streaming ST" must consume speech.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-speech-slide-probe.homorich-negara-gooya-pron-regen-audioinstructtts-three-model-gemini-zh
InstructTTSEval 三模型 Gemini 评测数据
本目录整理了 InstructTTSEval 中文集上三个 TTS 模型的生成音频和 Gemini 一致性评测结果:Qwen3-TTS-12Hz-1.7B-VoiceDesign、Seed-Audio-1.0、VoxCPM2。
字段
records.jsonl 每行对应一个模型和一种控制格式(APS、DSD 或 RP):
id:InstructTTSEval 样本 ID
mode:控制格式
model、model_name:模型标识
text:合成文本
instruction:历史评测记录中的输入控制指令,按本行 APS/DSD/RP 格式保留;不等同于各模型 API 的完整请求封装
generated_audio:该模型生成音频的相对路径
reference_audio:原始参考音频的相对路径
gemini_consistent:Gemini judge 的一致性判断
inconsistency_reason:判断为不一致时的原因… See the full description on the dataset page: https://huggingface.co/datasets/zsy814/instructtts-three-model-gemini-zh.OpenGameArt-GPL-2.0
Dataset Card for OpenGameArt-GPL-2.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 2.0 (GPL-2.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-GPL-2.0.OpenGameArt-GPL-3.0
Dataset Card for OpenGameArt-GPL-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-GPL-3.0.OpenGameArt-GPL-3.0
Dataset Card for OpenGameArt-GPL-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/irfankabir02/OpenGameArt-GPL-3.0.grocery-bench
Grocery Bench
30-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a grocery ordering assistant.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a grocery ordering assistant helping a customer build, modify, and finalize an order. The conversation is designed around 15 difficulty enhancements that… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/grocery-bench.goodforft
goodforft
This is a merged speech dataset containing 863 audio segments from 4 source datasets.
Dataset Information
Total Segments: 863
Speakers: 4
Languages: en
Emotions: angry, happy, neutral
Original Datasets: 4
Dataset Structure
Each example contains:
audio: Audio file (WAV format, original sampling rate preserved)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/goodforft.chinese-lips-longform-debug
Chinese-LiPS Long-Form (zh long streaming speech)
Reconstructed continuous long-speech streams from
BAAI/Chinese-LiPS, for
slide-aware / streaming speech-translation development and evaluation. Each
source video (one speaker, one scripted lecture with slides) was released as
pre-segmented clips; here they are re-joined into the full talk.
Two variants of the same 3 talks (~97 min speech total):
config
how segments are placed
use
orig_timeline
at their original session… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-longform-debug.Gurbani-MahanKosh-Frontier-Corpus
ੴ Gurbani & Bhai Kahn Singh Nabha Mahan Kosh Frontier Corpus
☬ ਗੁਰਬਾਣੀ ਅਤੇ ਭਾਈ ਕਾਹਨ ਸਿੰਘ ਨਾਭਾ 'ਮਹਾਨ ਕੋਸ਼' ਪ੍ਰਮਾਣਿਕ ਡਾਟਾਸੈੱਟ
👨💻 Project Lead & Architecture
Curator & Developer: Gurpreet Singh Dhillon (Nam-toon Studio)
GitHub Profile: github.com/gurpreetsingh5523-source
Flagship Project: AMRIT Research OS (Autonomous Medical AI)
📖 Dataset Overview
An authoritative lexical dataset compiling authentic definitions… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Gurbani-MahanKosh-Frontier-Corpus.MultiFraudAlign
FraudAlign-MCS
A fraud-only multilingual & code-switched dataset of scam-call dialogues, natively generated
(not translated) with Qwen2.5-72B-Instruct-AWQ. Modeled on the schema, fraud
taxonomy, and per-type proportions of the Chinese TeleAntiFraud-28k dataset,
regenerated from scratch in 4 languages: English (en), Hindi (hi), Korean (ko), Hinglish (Hindi-English code-switch) (hinglish).
28,708 dialogues total (7,177 per language), built to support
alignment of audio language… See the full description on the dataset page: https://huggingface.co/datasets/ggirishg/MultiFraudAlign.soundstock.com-music-genres-taxonomy
SoundStock Music Genres Taxonomy
A large, structured, and extensible music genre taxonomy dataset designed for music tagging, classification, search, recommendation systems, and audio / music machine learning workflows.
This dataset provides a hierarchical view of music genres, including root genres, subgenres, and expanded variants (style, era, region, and fusion), with stable IDs suitable for long-term use in production systems.
📊 Dataset Overview
1,600+ genres… See the full description on the dataset page: https://huggingface.co/datasets/SoundStock/soundstock.com-music-genres-taxonomy.gemma-4-e4b-audio-qa
Gemma-4 E4B Audio-QA Training Mix
A 91k-row audio question-answering dataset assembled from four public upstream
datasets, formatted as ChatML-style conversations for instruction-tuning an
audio-language model. This is the exact training data used for
bnovikov/gemma-4-e4b-audio-v3.
Important: this repository contains only the metadata and prompts/answers.
The audio files are NOT hosted here. Each audio_path is a source-tagged ID
like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.small-german-medical-dialogue-dataset-for-moshi
Small german dialogue dataset
This dataset contains 500 completely made up medical phonecall dialogues between patients and a GP's office.
Dataset Details
Dataset Description
500 made up phonecalls that were first created with AI as text.
The audio was then created using Openai tts-1-hd and the accurately timestamped transcripts were added.
The audio files are formatted like this:
Stereo with split channels:
Speaker A is on the left channel… See the full description on the dataset page: https://huggingface.co/datasets/chtugha/small-german-medical-dialogue-dataset-for-moshi.odyssey-v0.5.1-16h-eval-audio-gated
🔒 Odyssey V0.5.1 — 16H Gated Audio Evaluation Repo
This repository contains the gated audio companion to the public Odyssey evaluation preview.
All requests are reviewed manually.
Public Preview Repository
For transcripts, schema preview, and metadata inspection, see:
ODYSSEYAILABS/odyssey-v0.5.1-16h-eval-preview-public
About Odyssey AI Labs
Odyssey AI Labs builds premium African data infrastructure for next-generation AI systems.
chlid-datasetGenshin4.8_JPhomorich-negara-gooya-grapheme-regen-audioGawrGuraua_ru_mix_citrinet_512_gamma_0_25medimind-r11-train
MediMind R11 — ASR training data
Unified manifest + packed audio for fine-tuning Whisper-large-v3 on Norwegian
clinical and conversational speech.
Training manifest: r11_manifest.jsonl — 11,022 packs
Held-out eval set: r11_heldout_eval.jsonl — 291 packs (NEVER train on these)
~see manifest audit packs total
11 sources: lege_*, podcasts (motiv/podk/stet), nb_samtale, nb_tale_m3, tts_drugs
Schema
See r11_manifest.jsonl (one JSON object per line) and… See the full description on the dataset page: https://huggingface.co/datasets/gallip0li/medimind-r11-train.
