datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mls_sidon
MLS-Sidon
Overview
This dataset is a cleansed version of Multilingual LibriSpeech (MLS) with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling.
The dataset is provided in WebDataset format for efficient large-scale training.
Source: Multilingual LibriSpeech
Languages: English, German, French, Spanish, Italian, Polish, Dutch, Portuguese
Format: WebDataset (.tar shards)
License: CC-BY-4.0
Dataset Structure
Each sample in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/mls_sidon.audiosnippetsClotho-Moment
Clotho-Moment
This repository provides wav files used in Language-based Audio Moment Retrieval.
Each sample includes long audio containing some audio events with the temporal and textual annotation.
Project page: https://h-munakata.github.io/Language-based-Audio-Moment-Retrieval/
Code: https://github.com/line/lighthouse
Split
Train
train/train-{000..715}.tar
37930 audio samples
Valid
valid/valid-{000..108}.tar
5741 audio samples
Test
test/test-{000..142}.tar
7569… See the full description on the dataset page: https://huggingface.co/datasets/lighthouse-emnlp2024/Clotho-Moment.audiosnippets_small_with_detailed_annotationaudiosnippets_small_with_detailed_annotation2captioned-ai-music-snippets
Dataset Overview
A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models.
Source
Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository.
Captioning
All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions.
License
Apache 2.0
mls_hq_urgent_track1majestrino-unified-detailed-captions
Majestrino Unified Detailed Captions
Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption.
Stats
4,658,407 samples
932 tar files (~1.1 GB each)
~1,017 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption
caption_type — always unified_detailed_caption
transcription — speech transcription (when available, normalized from multiple source keys)
duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.majestrino-dataaudiosnippets_long_2_5Mlaions_got_talent_enhanced_no_metadataai-maudiosnippets_long_1Mai-music-deduplicated
AI Music Deduplicated
A large-scale collection of AI-generated music from five platforms: Mureka, Riffusion, Sonauto, Suno, and Udio. Each song includes the original audio file and its full platform metadata as a JSON sidecar.
Overview
Subset
Songs
Tar Files
Size
Audio Format
Source Platform
mureka
~312K
49
~981 GB
.mp3
Mureka
riffusion
~105K
14
~266 GB
.m4a
Riffusion
sonauto
~15K
2
~25 GB
.ogg
Sonauto
suno
~307K
65
~1.3 TB
.mp3
Suno
udio~126K
33
~642… See the full description on the dataset page: https://huggingface.co/datasets/ai-music/ai-music-deduplicated.audiosnippets-cleaned
Dataset Summary
This dataset is a processed version of mitermix/audiosnippets. The dataset contains audio snippets that have been cleaned and resampled, making it suitable for tasks like audio captioning, audio classification, or other audio-based machine learning applications.
Processing Details
Transcriptions and broken characters were removed.
All MP3 audio files were resampled to 16kHz for consistency.
The accompanying JSON metadata was made consistent.
Entries with… See the full description on the dataset page: https://huggingface.co/datasets/mkrausio/audiosnippets-cleaned.qwen3-tts-multilingual-emotional-speechMemeEffect-382K-audioWe are releasing the audio files that we have collected from MemeEffect-382K dataset. All the files are being shared as .tar files and files are rnamed using their respective id that can be found through the metadata.
We share these files as-part of research initiative.
MECAT-QAMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-Caption (HF)
Dataset Description
MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks:
Audio Captioning: Generating textual descriptions for given audio
Audio Question Answering: Answering questions about given audio… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-QA.multi_round_speech_180kMMAudio-precomputed-results
Precomputed results for MMAudio
Results from four model variants of MMAudio.
All results are in the .flac format with lossless compression.
A cache folder contains the feature caches computed by the evaluation script.
Code: https://github.com/hkchengrex/MMAudio
Evaluation: https://github.com/hkchengrex/av-benchmark
VGGSound
Contains the VGGSound test set results. There are 15220 videos, collected with our best effort. Not all videos in the test sets are available… See the full description on the dataset page: https://huggingface.co/datasets/hkchengrex/MMAudio-precomputed-results.german_mmusan-mirrorxares_llm_dataMECAT-CaptionMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-QA (HF)
Dataset Description
MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks:
Audio Captioning: Generating textual descriptions for given audio
Audio Question Answering: Answering questions about given audio
Generated via… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-Caption.musanMUSAN
Identifier: SLR17
Summary: A corpus of music, speech, and noise
Category: Audio
License: Attribution 4.0 International (CC BY 4.0)
Downloads (use a mirror closer to you):
musan.tar.gz [11G] ( The corpus ) Mirrors: [EU] [EU] [CN]
About this resource:
MUSAN is a corpus of music, speech, and noise recordings.
This work was supported by the National Science Foundation Graduate Research Fellowship under Grant No. 1232825 and by Spoken Communications.
You can cite the data using the… See the full description on the dataset page: https://huggingface.co/datasets/EaseZh/musan.majestrino-1.00-16xk5-sae-features
Majestrino 1.00 SAE — Feature Audio Samples (16x, k=5)
Top-2000 activating audio samples for each feature in the
Majestrino 1.00 SAE.
Overview
Metric
Value
SAE Architecture
16x expansion, k=5, d_model=768
Total Features
12,288
Alive Features
10,684
Audio per Feature
Up to 2,000 highest-activating
Audio Format
Opus (24 kbps OGG container)
Total TAR Files
1069
Source Dataset
laion/majestrino-data
File Structure
Each TAR file… See the full description on the dataset page: https://huggingface.co/datasets/laion/majestrino-1.00-16xk5-sae-features.nug_myanmar_asr
366 Hours NUG Myanmar ASR Dataset
The NUG Myanmar ASR Dataset is the first large-scale open Burmese speech dataset — now expanded to over 521,476 audio-text pairs, totaling ~366 hours of clean, segmented audio. All data was collected from public-service educational broadcasts by the National Unity Government (NUG) of Myanmar and the FOEIM Academy.
This dataset is released under a CC0 1.0 Universal license — fully open and public domain. No attribution required.
🕊️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/nug_myanmar_asr.mls-enhanced-dacvae
Multilingual LibriSpeech converted to DAC VAE latents
Source
facebook/multilingual_librispeech
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-enhanced-dacvae.voa_myanmar_asr_audio_1
📢 This is the first publicly released ASR-ready Burmese speech dataset with over 1 million audio chunks — a milestone in the history of Myanmar language technology.
Overview
This dataset was created by scraping and segmenting the full archive of the VOA Burmese morning radio program. Out of a total of 3,687 full-length MP3 broadcasts, this release processes 3,267 of them, resulting in approximately 1.8 million sentence-level audio chunks, totaling ~3,267 hours of segmented audio.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_1.MM-PreTrain
JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
[HomePage]
[Paper]
[GitHub]
TL;DR
We introduce JavisGPT, a multimodal LLM that can understand audiovisual inputs and simultaneously generate synchronized sounding videos in a unified model.
We also curate the JavisInst-Omni dataset to facilitate instruction-tuning for comprehension and generation on sounding videos.
📰 News
[2025.12.30] 🚀 We release the training… See the full description on the dataset page: https://huggingface.co/datasets/JavisVerse/MM-PreTrain.
