datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
otoSpeech-full-duplex-turn-104h
Dataset Card for otoSpeech-full-duplex-turn-104h
Contact
Website: https://oto.earthEmail: agent@oto.earth
Dataset Summary
otoSpeech-full-duplex-turn-104h is an English, full-duplex conversational speech dataset for research on turn-taking and related spoken-dialogue phenomena. It contains 420 two-speaker conversations totaling approximately 104.94 hours. Each conversation includes time-aligned, channel-separated audio, a stereo combined recording… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-turn-104h.free-music-archive-full
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-full.freesound-laion-640k-commercial-16khz-full
About this Repository
This repository is the training split of the complete FreeSound LAION 640k dataset, limited only to licenses that permit commercial works, resampled to 16khz using torchaudio.transforms.Resample.
This is ideal for use cases where a variety of audio is desired but fidelity and labels are unnecessary, such as background audio for augmenting other datasets.
Dataset Versions
You are looking at the full dataset which contains 403,146 unique sounds… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/freesound-laion-640k-commercial-16khz-full.otoSpeech-full-duplex-280h
📢 Notice: We released a new processed version of otoSpeech. Available here: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-processed-141h
Dataset Card for otoSpeech-full-duplex-280h: Full-Duplex Conversational Speech Dataset
Contact
Website: https://oto.earth
Mail: consome@oto.earth
Dataset Summary
otoSpeech-full-duplex-280h is a 280-hour, full-duplex, two-speaker conversational speech dataset.
Each sample includes 48 kHz… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-280h.dataset_esc50_full_v1audioset-fulldataset_gise_full_v1tadabur-lora-data-fullfleurs-full
FLEURS Full - Test Set for ASR Benchmarking
Complete test set of Google FLEURS for all 30 languages supported by Qwen3-ASR, prepared for benchmarking with FluidAudio.
Languages (30)
Asian Languages (13)
Code
Language
Samples
cmn_hans_cn
Chinese (Mandarin)
945
yue_hant_hk
Cantonese
819
ja_jp
Japanese
650
ko_kr
Korean
382
vi_vn
Vietnamese
857
th_th
Thai
1,021
id_id
Indonesian
687
ms_my
Malay
749
hi_in
Hindi
418
ar_eg
Arabic (Egyptian)… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/fleurs-full.dataset_urbansound8k_full_v1MoulSot-Full
MoulSot-Full Dataset
Dataset Summary
MoulSot-Full is a large-scale Moroccan Darija speech dataset containing in total 1,500 hours of speech audio. From this extensive corpus, a high-quality subset of approximately 80 hours has been carefully curated and transcribed. It was built entirely from publicly available YouTube content across 51 diverse channels (including vlogs, podcasts, interviews, and commentary) to capture real-world Moroccan Darija, including natural… See the full description on the dataset page: https://huggingface.co/datasets/atlasia/MoulSot-Full.free-music-archive-commercial-16khz-full
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-commercial-16khz-full.2026-04-01T22-36-32plus00-00_gdpval_full220TV-44kHz-Full
The "Thorsten-Voice" dataset
This truly open source (CC0 license) german (🇩🇪) voice dataset contains about 40 hours of transcribed voice recordings by Thorsten Müller,
a single male, native speaker in over 38.000 wave files.
Mono
Samplerate: 44.100Hz
Trimmed silence at begin/end
Denoised
Normalized to -24dB
Disclaimer
"Please keep in mind, I am not a professional speaker, just an open source speech technology enthusiast who donates his voice. I contribute my personal… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-44kHz-Full.otoSpeech-full-duplex-processed-141h
Dataset Card for otoSpeech-full-duplex-processed-141h: Full-Duplex Conversational Speech Dataset
Dataset Summary
otoSpeech-full-duplex-processed-141h is a full-duplex, two-speaker conversational speech dataset. It is derived from otoSpeech-full-duplex-280h and has been curated and processed as follows:
Selected high-quality conversations based on human reviews.
Applied noise reduction and speech enhancement to improve audio quality.
Added new samples collected after the… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-processed-141h.pd-voice-full-multimodal-dataset
Parkinson Voice — Full Multimodal Dataset
Complete Parkinson’s vs healthy voice package for classification and explainable Mel reasoning research (EDGE).
Not Mel-only: raw audio, 10 visual modalities, feature CSVs, plus Gemma reasoning traces for Mel.
Contents
Path
Description
audio/
Waveform clips (healthy / parkinsons), 1134 files
images/mel/
Mel spectrograms
images/spectrogram/
Linear spectrograms
images/mfcc/
MFCC maps
images/delta_mfcc/… See the full description on the dataset page: https://huggingface.co/datasets/mdimamhosen/pd-voice-full-multimodal-dataset.malagasy-speech-fullFull-Duplex-Bench-Datavctk-fullBenglaAI_Full_Datasetsmartdj_fullMoulSot-Full
MoulSot-Full Dataset
Dataset Summary
MoulSot-Full is a large-scale Moroccan Darija speech dataset containing in total 1,500 hours of speech audio. From this extensive corpus, a high-quality subset of approximately 80 hours has been carefully curated and transcribed. It was built entirely from publicly available YouTube content across 51 diverse channels (including vlogs, podcasts, interviews, and commentary) to capture real-world Moroccan Darija, including natural… See the full description on the dataset page: https://huggingface.co/datasets/MathematicianNLPer/MoulSot-Full.anv-data-ke-somali-fullanv-data-ke-somali-fullnavigation-corpus-speech-full-dagbani
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana TTS Navigation Corpus — Dagbani
Synthetic speech dataset for navigation.
Structure
audio/ – all .wav audio files
text/ – matching .txt files with transcriptions
metadata.csv – full metadata table
Omni-DuplexEval-Full
Omni-DuplexEval
Omni-DuplexEval is a benchmark for evaluating real-time duplex multimodal interaction. Unlike conventional offline video understanding benchmarks, Omni-DuplexEval focuses on streaming settings where models must continuously process evolving multimodal inputs and decide what to respond and when to respond.
The benchmark contains two scenarios:
Real-Time Description (RTD): evaluates continuous streaming description ability.
Proactive Reminder (PR): evaluates… See the full description on the dataset page: https://huggingface.co/datasets/tianrantianran/Omni-DuplexEval-Full.watkins-marine-mammal-full-cuts
Watkins Marine Mammal Sound Database
Dataset Description
The Watkins Marine Mammal Sound Database (WMMSD) is one of the largest historical collections of marine mammal vocalizations. It contains 15,248 recordings spanning nearly seven decades from 54 marine mammal species, including whales, dolphins, porpoises, seals, sea lions, manatees, sea otters, and other marine mammals.
This repository provides the complete dataset in a format fully compatible with the… See the full description on the dataset page: https://huggingface.co/datasets/ivangtorre/watkins-marine-mammal-full-cuts.kalenjin-speech-fullmusiccaps_processed_full
Dataset Card for "musiccaps_processed_full"
This dataset is only for educational and research purposes.
All the licenses valid for google/MusicCaps dataset and its content apply here as well.
More Information needed
video-full-duplex-benchmark
VideoFDB: Video-Full-Duplex-Benchmark
Project Page · HuggingFace · Paper (arXiv)
Dataset Description
A benchmark dataset of annotated, two-person video conference recordings designed to support the evaluation of multimodal AI agents in conversational settings. The dataset covers 11 distinct conversational dynamics — spanning verbal, nonverbal, and mixed-modality behavior — annotated through a three-pass human-in-the-loop pipeline.
The benchmark consists of trimmed… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/video-full-duplex-benchmark.
