datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
S2SBench
S2SBench
📄 View Paper 📥 Code
S2SBench is a benchmark designed to evaluate the intelligence degradation of speech-to-speech large language models.
The Dataset
S2SBench includes three evaluation sets:
sStoryCloze: English speech-based story cloze task.
zh-sStoryCloze: Chinese speech-based story cloze task.
sCMMLU: Speech-based version of CMMLU, covering multiple-choice questions across various disciplines.
Dataset Statistics
Dataset
Sample Pairs… See the full description on the dataset page: https://huggingface.co/datasets/undobug/S2SBench.thomcles-persian-farsi-speech-whisper-segmented-under30suts2025_vietipa
Vietnamese IPA Dataset
A comprehensive Vietnamese IPA (International Phonetic Alphabet) dataset with word pronunciations and MP3 audio files for text-to-speech and pronunciation learning applications.
Dataset Description
Dataset Summary
This dataset contains 50 common Vietnamese words with their IPA (International Phonetic Alphabet) transcriptions and corresponding audio files. It's designed for:
Text-to-speech systems development
Vietnamese pronunciation… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/uts2025_vietipa.librispeech-asr-whisper-segmented-under30snorthtts-men-v3-whisper-segmented-under30sfleurs-farsi-whisper-segmented-under30smodified-shemo-whisper-segmented-under30sneyshekar-v6-hf-whisper-segmented-under30sAudio-Understanding-Test-Set
Audio Understanding Test Set
A structured dataset for evaluating audio understanding capabilities of multimodal AI models. Contains 137 test prompts across 22 categories, paired with a 20-minute voice sample and 49 completed model outputs from Gemini 3.1 Flash Lite.
Overview
Property
Value
Total prompts
137
Implemented (with prompt text)
49
Suggested (description only)
88
Completed outputs
49
Categories
22
Model under test… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Audio-Understanding-Test-Set.Underwater_Audio
Watkins Marine Mammal Sound (WMMS) Database
Sound files on this website are free to download for personal or academic (not commercial) use.
Sound files and associated metadata are credited as follows: "Watkins Marine Mammal Sound Database, Woods Hole Oceanographic Institution and the New Bedford Whaling Museum."
Database could be found and downloaded from here.
In this database version, the audio archive includes sounds of 32 species:
Atlantic_Spotted_Dolphin
Bearded_Seal
Beluga… See the full description on the dataset page: https://huggingface.co/datasets/sangkrishna/Underwater_Audio.demoMr.Underhill_Vtuberunderstand_only_wavstillalive-overlap-rule-not-flagged-under2s-1000-20260922audio-understandingtest_data_FWstillalive-overlap-rule-flagged-under2s-1000-20260922tesssunder_30_male_part_3sbpn-full-music-snr-difference-under-2db-20260722
Full music recordings with Librosa SNR difference below 2 dB
This manually gated dataset contains 4 complete recordings. Every
included recording had at least one transcript chunk labelled as background
music, while its full-recording Librosa score satisfied:
snr_after_db - snr_before_db < 2.0
snr_before_db is measured on the complete original recording and
snr_after_db is measured on the complete Demucs vocals recording. The
audio column contains the original full recording… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/sbpn-full-music-snr-difference-under-2db-20260722.sanskrit_audio_dataset_under_30audio-understanding-dataUndea2
