datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
youtube_caption_yue
YouTube ASR Caption Dataset (Cantonese)
This dataset was built from YouTube videos with manually provided captions in Cantonese. We used SenseVoice to re-transcribe the audio and filtered segments to build a high-quality collection of audio-caption pairs.
What’s included
Segments where the ASR output is identical to the original caption — likely clean.
Segments where differences are only homophones (同音字) or English words — likely ASR mistakes.
This combination supports… See the full description on the dataset page: https://huggingface.co/datasets/ming030890/youtube_caption_yue.viVoice
Important Note ⚠️
This dataset is only to be used for research purposes. Access requests must be made via your school, institution, or work email. Requests from common email services will be rejected. We apologize for any inconvenience.
viVoice: Enabling Vietnamese Multi-Speaker Speech Synthesis
For a comprehensive description, please visit https://github.com/thinhlpg/viVoice
This dataset is licensed under CC-BY-NC-SA-4.0 and is intended for research purposes only.… See the full description on the dataset page: https://huggingface.co/datasets/capleaf/viVoice.captioned-ai-music-snippets
Dataset Overview
A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models.
Source
Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository.
Captioning
All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions.
License
Apache 2.0
majestrino-unified-detailed-captions
Majestrino Unified Detailed Captions
Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption.
Stats
4,658,407 samples
932 tar files (~1.1 GB each)
~1,017 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption
caption_type — always unified_detailed_caption
transcription — speech transcription (when available, normalized from multiple source keys)
duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.egocentric-vr-capture-20h-multimodal-sample
Egocentric VR Capture — 20-Hour Multimodal Inspection Sample
195 real-world task episodes / 2,283,482 frames / 21.14 delivered hours captured with consumer VR hardware. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package.
This publicly accessible 20-hour-scale dataset is produced by the EXYLOS real-world data pipeline. Files and the Dataset Viewer can be accessed without individual approval;… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/egocentric-vr-capture-20h-multimodal-sample.capes_synthetic_audio_filteredaudioset_captioncc0-music-captionedCollected from various CC0 music sites. These are all instrumental - none have lyrics. They are annotated with a description of the song.
Some of the music comes from FreePD, a site that shared public domain music. The FreePD website has since been taken down.
All songs were created by humans, not AI-generated.
Closed_Captioning_Lecture_DatasetCapTTS-SFT-ears-cleanedMECAT-CaptionMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-QA (HF)
Dataset Description
MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks:
Audio Captioning: Generating textual descriptions for given audio
Audio Question Answering: Answering questions about given audio
Generated via… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-Caption.tacos-captioningcommon_voice_22_yue_w_background_captionMerged JackyHoCL/urban-noise-uganda-61k-caption, OpenSound/AudioCaps
TODO: convert to MP3, reduce size
gurbani-sehajpath-yt-captions-canonical
Gurbani Sehajpath — Canonical-aligned ASR corpus
Stage-1 + Stage-2 canonical pipeline output for sehaj-path (calm recitation of the Guru Granth Sahib). Built from publicly available audio recordings with aligned transcripts, chunked by caption timing and aligned against the canonical Guru Granth Sahib Ji text (SGGS).
Columns
Schema is auto-inferred from the parquet shards. Primary columns:
audio — 16 kHz mono waveform
final_text — canonical Gurmukhi transcription (post… See the full description on the dataset page: https://huggingface.co/datasets/surindersinghssj/gurbani-sehajpath-yt-captions-canonical.egocentric-vr-capture-1h-multimodal-sample
Egocentric VR Capture — 1-Hour Multimodal Inspection Sample
13 real-world task episodes / 108,029 frames / approximately 60 minutes captured with Meta Quest 3. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package.
This publicly accessible dataset is an inspection slice produced by the EXYLOS real-world data pipeline. It demonstrates capture quality, synchronization, schema, and QA metadata… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz042/egocentric-vr-capture-1h-multimodal-sample.spectrogram-captionsDataset of captioned spectrograms (text describing the sound).
timbre-audio-caption-pairsfreesound-commercially-permissive-subset-with-captionscaptionercv22_yue_captionaudioset-with-grounded-captions1125_trad_sample_16k_caption_RoformerSeparateoperation-cappicinoXCT_Data_of_Retired_Large_Capacity_Prismatic_BatteriesThis dataset provides 3D XCT reconstruction results for eight retired and one fresh LiFePO4 (LFP) prismatic batteries. These samples were manufactured by REPT BATTERO Energy Co., Ltd. in 2021 and sourced from Dongfeng E70 electric vehicles. Each battery has a nominal voltage of 3.2 V and a rated capacity of 135 Ah (dimensions: 148 mm × 80 mm × 105 mm). Compared to the fresh control sample (136.7 Ah), the retired batteries exhibit varying degrees of degradation, with measured capacities ranging… See the full description on the dataset page: https://huggingface.co/datasets/JiaWeiDong/XCT_Data_of_Retired_Large_Capacity_Prismatic_Batteries.music_caps_4sec_wave_typeemo_speech_caption_testbinary-classifier-birdnet
Binary BirdNet Classifier
Contiene anotaciones y audios de 3s y 5s para clasificación binaria con rutas relativas.
music-audio-pseudo-captions
Dataset Card for Music-Audio-Pseudo Captions
Pseudo Music and Audio Captions from LP-MusicCaps, Music Negation/Temporal Ordering WavCaps
Dataset Summary
Compared to other domains, music and audio domains cannot obtain well-written web caption data, and caption annotation is expensive.
Therefore, we use the Music (LP-MusicCaps), (Music Negation/Temporal Ordering) and Audio (Wavcaps) datasets created with ChatGPT to re-organize them in the form of instructions, input… See the full description on the dataset page: https://huggingface.co/datasets/seungheondoh/music-audio-pseudo-captions.majestrino-unified-detailed-captions-temporal
Majestrino Unified Detailed Captions with Temporal Aspects
Filtered subset of laion/majestrino-data containing only samples with unified_detailed_caption_with_temporal_aspects.
Stats
4,128,665 samples
826 tar files (~1.1 GB each)
~878 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption with temporal aspects
caption_type — always unified_detailed_caption_with_temporal_aspects
transcription — speech… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions-temporal.audioset-with-captions
