CoolFace
Datasetpublic

AxonData/background-noise-detection-dataset

Speech-Free Background Noise Dataset — Real-World, Non-Synthetic (50+ Hours) Dataset summary 50+ hours of real-world urban environmental/ambient background noise (field recordings) without intelligible speech (speech-free), from three scenes: airport, street, subway. The dataset is non-synthetic and intended for speech enhancement via noise augmentation and sound event detection (SED) as “clean background”/negative class Full version of dataset is… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/background-noise-detection-dataset.

sourceHugging Facecc-by-nc-4.0updated 7mo agoView on Hugging Face
1likes17downloads
Dataset Card

Speech-Free Background Noise Dataset — Real-World, Non-Synthetic (50+ Hours)

Dataset summary

50+ hours of real-world urban environmental/ambient background noise (field recordings) without intelligible speech (speech-free), from three scenes: airport, street, subway. The dataset is non-synthetic and intended for speech enhancement via noise augmentation and sound event detection (SED) as “clean background”/negative class

Full version of dataset is availible for commercial usage - leave a request on our website Axonlabs to purchase the dataset 💰

Purpose and usage scenarios

  • —Speech enhancement: adding ambient/background noise to clean speech
  • —Sound Event Detection: background samples without target events/speech; negative samples and false alarm rate estimation
  • —Filtering/noise reduction: training noise reduction models without the risk of intelligible speech leakage

Noise Environments

  • —airport: terminals, corridors, gates, baggage areas — ambient/background noise
  • —street: sidewalks and roadways; traffic, wind, footsteps, street music as indistinct background ambient noise
  • —subway: platform, train car, passageways; braking/acceleration, tunnel rumble, doors, announcements as indistinct background noise

Features:

  • —Only real-world field recordings. No synthetic mixes; non-synthetic source audio
  • —No intelligible speech (speech-free). Natural crowd murmur allowed only when no single utterance is intelligible
  • —Only noise. Music, dominant speech, and close-up announcements are excluded