datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ears
EARS: Expressive Anechoic Recordings of Speech
This is a mirror of the Expressive Anechoic Recordings of Speech (EARS) dataset.
The original files were converted from WAV to Opus to reduce the size and accelerate streaming.
Sampling rate: 48 kHz
Channels: 1
Format: Opus
Splits:
Train: 92 hours, 15939 utterances, speakers p001 to p099
Validation: 2 hours, 322 utterances, speakers p100 and p101
Test: 6 hours, 966 utterances, speakers p102 to p107
License: CC BY-NC 4.0
Source:… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/ears.fsd50k
FSD50K: An open dataset of human-labeled sound events
This is a mirror of the FSD50K sound event dataset.
The original files were converted from WAV to Opus to reduce the size and accelerate streaming.
Sampling rate: 48 kHz
Channels: 1
Format: Opus
Splits:
Dev: 80 hours, 40966 clips.
Eval: 28 hours, 10231 clips.
License: FSD50K is released under CC-BY. However, each clip has its own licence. Clip licenses include CC0, CC-BY, CC-BY-NC and CC Sampling+. Clip licenses are specified… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/fsd50k.wham
WHAM!48kHz noise dataset
This is a mirror of the WHAM!48kHz noise dataset.
The original files were segmented and converted from WAV to Opus to reduce the size and accelerate streaming.
Sampling rate: 48 kHz
Channels: 2
Format: Opus
Splits:
Train: 59 hours, 21216 segments, files 000 to 188
Validation: 12 hours, 4444 segments, files 189 to 225
Test: 7 hours, 2613 segments, files 226 to 249
License: CC BY-NC 4.0
Source: http://wham.whisper.ai/
Paper: WHAM!: Extending Speech… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/wham.dns5
DNS5 Challenge data
This is a mirror of the DNS5 Challenge data.
The original files were converted from WAV to Opus to reduce the size and accelerate streaming.
⚠️ Only the LibriVox, AudioSet, Freesound, OpenSLR26, and OpenSLR28 data is included. The VCTK, VocalSet, CREMA-D, VoxCeleb2, and DEMAND data is excluded. ⚠️
Sampling rate: 48 kHz
Channels: 1
Format: Opus
Splits:
speech_english: 245 hours, 186743 files
speech_french: 95 hours, 60454 files
speech_german: 137 hours, 119175… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/dns5.libri
LibriSpeech: An ASR corpus based on public domain audio books
This is a mirror of the LibriSpeech ASR corpus.
The original files were converted from FLAC to Opus to reduce the size and accelerate streaming.
The transcripts are not included. This mirror is thus best suited for audio-to-audio tasks.
Sampling rate: 16 kHz
Channels: 1
Format: Opus
Splits:
Train: 460 hours, 132553 utterances, train-clean-100 and train-clean-360 sets.
Validation: 7 hours, 2703 utterances, dev-clean set.… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/libri.vctk
VCTK
This is a mirror of the VCTK Corpus.
The original files were converted from FLAC to Opus to reduce the size and accelerate streaming.
Sampling rate: 48 kHz
Channels: 1
Format: Opus
Splits:
train_mic1: 90 speakers, 33.6 hours, 35987 utterances
train_mic2: 90 speakers, 33.6 hours, 35987 utterances
val_mic1: 10 random speakers unseen during training: p238, p244, p254, p263, p265, p272, p288, p294, p305, and p335. 4.0 hours, 4179 utterances.
val_mic2: Same speakers as val_mic1.… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/vctk.mls-hq-urgent-track1
Multilingual LibriSpeech HQ (MLS-HQ)
This is a mirror of the Multilingual LibriSpeech HQ (MLS-HQ) data used in URGENT 2025 Track 1.
The original files were converted from FLAC to Opus to reduce the size and accelerate streaming.
Sampling rate: 48 kHz (resampled from 44.1 kHz to support Opus format)
Channels: 1
Format: Opus
Splits:
spanish: 150 hours, 36031 utterances
german: 150 hours, 35890 utterances
french: 150 hours, 36078 utterances
License: CC0 1.0
Source:… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/mls-hq-urgent-track1.fma-medium
Free Music Archive (FMA-medium)
This is a mirror of FMA-medium.
Sampling rate: 24 and 48 kHz
Channels: 1 and 2
Format: Opus
Duration: 208 hours, 24908 tracks
License:
Each track is distributed under the license chosen by the artist. See tracks.csv for details.
The metadata is distributed under CC BY 4.0.
Source: https://github.com/mdeff/fma
Paper: FMA: A Dataset For Music Analysis
Usage
import io
importsoundfile as sf
from datasets import Features, Value… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/fma-medium.demand
DEMAND: Diverse Environments Multichannel Acoustic Noise Database
This is a mirror of the DEMAND: Diverse Environments Multichannel Acoustic Noise Database.
The original files were segmented and converted from WAV to Opus to reduce the size and accelerate streaming.
Sampling rate: 16 kHz and 48 kHz
Channels: 1 (individual channels are distributed as separate files)
Format: Opus
Splits:
16k: 22.7 hours, 6800 segments (each channel is treated as a separate file)
48k: 24.0 hours, 7200… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/demand.tut2020
TUT Urban Acoustic Scenes 2020 Mobile Development and Evaluation datasets
This is a mirror of the TUT Urban Acoustic Scenes 2020 Mobile Development and Evaluation datasets.
The original files were converted from WAV to Opus to reduce the size and accelerate streaming.
Sampling rate: 48 kHz (resampled from 44.1 kHz to support Opus format)
Channels: 1
Format: Opus
Splits:
Dev: 64 hours, 23035 clips.
Eval: 33 hours, 11880 clips.
License: Custom. See LICENSE file.
Source:… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/tut2020.kenya-philippines-twospeaker-english-dialogue
Kenya/Philippines English Dialogue
Two-speaker dialogues in English, recorded on split tracks.
Changelog
Jan 2026: v1 release - vad-segmented WebRTC tracks
Specs
Speakers: >150; ~15 PH, remaining KE
Total duration: ~65 hours
Files sample rate: 48kHz
Actual sample rate: TBD
Language: English (PH, KE accents)
Topics: day-to-day conversation
Collection method
The dataset is built to capture the variety in the Kenyan accent.
The Philippino interviewers… See the full description on the dataset page: https://huggingface.co/datasets/Reord-AI/kenya-philippines-twospeaker-english-dialogue.423-Hours-English-Philippine-Full-Duplex-Multi-Channel-Speech-Dataset
Description
필리핀 영어(English-Philippine) 멀티스트림 자연 대화 스마트폰 음성 데이터셋으로, 주어진 주제를 바탕으로 대화하는 방식으로 수집되었습니다. 전사 텍스트와 함께 화자 ID, 성별, 연령 등의 속성 정보가 제공됩니다. 다양한 지역의 폭넓은 화자로부터 데이터를 수집하여 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상을 지원합니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1771?source=hf.kr
Specifications
Format
16kHz, 16 bit, 비압축 WAV, 모노 채널, 화자별 채널 분리
Content category
주어진 주제를 기반으로 대화
Recording condition
낮은… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/423-Hours-English-Philippine-Full-Duplex-Multi-Channel-Speech-Dataset.IntentTrain
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
This repository contains the dataset and associated information for the paper HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context.
👀 HumanOmniV2 Overview
With the rapid evolution of multimodal large language models, the capacity to deeply understand and interpret human intentions has emerged as a critical capability, which demands detailed and thoughtful reasoning. In recent studies… See the full description on the dataset page: https://huggingface.co/datasets/PhilipC/IntentTrain.TIMIT_datasetclarity
Clarity Speech Corpus
This is a mirror of the Clarity Speech Corpus.
The original files were converted from WAV to Opus to reduce the size and accelerate streaming.
Sampling rate: 48 kHz (resampled from 44.1 kHz to support Opus format)
Channels: 1
Format: Opus
Duration: 9 hours, 11352 utterances
License: CC BY 4.0
Source: https://doi.org/10.17866/rd.salford.16918180
Paper: Dataset of British English speech recordings for psychoacoustics and speech processing research: The Clarity… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/clarity.600-Hours-English-Philippine-Full-Duplex-Multi-Channel-Speech-Data-Sample
Description
English(Philippine) Multi-stream Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks.
For more details, please refer to the link: https://www.nexdata.ai/datasets/speechrecog/1771?source=Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/600-Hours-English-Philippine-Full-Duplex-Multi-Channel-Speech-Data-Sample.drone-audio-detection-samples
Dataset Description
Drone Audio Detection Samples (DADS) is currently the largest publicly available drone audio database, specifically designed for developing drone detection systems using deep learning techniques. All audio files are standardized to a sample rate of 16,000 Hz, 16-bit depth, mono-channel, and vary in length from 500 milliseconds to several minutes.
Most drone audio files were manually trimmed to ensure that a drone was always present in the recording. However… See the full description on the dataset page: https://huggingface.co/datasets/phila415/drone-audio-detection-samples.audio-file-samplesvoxpopuli_es-ja
Dataset Card for Spanish-to-Japanese Automatic Speech Recognition Dataset
Dataset Summary
This dataset was created as part of a workshop organized by Yasmin Moslem, focusing on speech-to-text pipelines.
The workshop's primary goal is to enable accurate transcription and translation of spoken a source language into a written language (and learn how to do so, of course 😃)
The dataset serves as the foundation for developing and evaluating various models, including… See the full description on the dataset page: https://huggingface.co/datasets/Philou134/voxpopuli_es-ja.Philip
Podcast Stt Data
Available Subsets
This dataset contains the following video transcriptions:
Subset ID
Broadcaster
Parquet File
FRTpI2Gu1KA
BeerBiceps
FRTpI2Gu1KA_BeerBiceps/train.parquet
Generated automatically by BG Remover Data Maker
AudibleLight_Eigenmike32-5_DCASE-STARSS23_Dataset
AudibleLight Eigenmike32-5 DCASE-STARSS23 Dataset
This dataset contains 121 synthetic spatial audio scenes — 111 for training and 10 for evaluation — of 60 seconds each, generated with the AudibleLight dataset generator (DOI). Each scene is rendered as five independent simulated Eigenmike32 captures, with 32 channels per capture, resulting in 570 minutes of multichannel audio in total at 24 kHz.
Foreground Audio
Foreground events are sampled from ESC-50: Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/PhilippXXY/AudibleLight_Eigenmike32-5_DCASE-STARSS23_Dataset.phi-asr-extendedmedical_whisperphi4_audiophilbrrugratscresvoiceLupitarbd8updated_philippineLupitarbd3voice_Philippa_Perry
