background noise
CAIMAN-ASR-BackgroundNoise
Dataset Card for Myrtle/CAIMAN-ASR-BackgroundNoise
This dataset provides background noise audio, suitable for noise augmentation
while training Myrtle.ai's CAIMAN-ASR models.
Dataset Details
Dataset Description
Curated by: Myrtle.ai
License: Myrtle.ai's modifications to the source data are licensed under
the CC BY 4.0 license.
Some of the original data is under the CC BY 3.0 license; the rest is in the public domain.
Please see the Source Data section… See the full description on the dataset page: https://huggingface.co/datasets/Myrtle/CAIMAN-ASR-BackgroundNoise.Internet-background-noise
Internet Background Noise Dataset (Unlabeled Raw Data)
This dataset contains HTTP internet noise data collected by an internet honeypot. It consists of raw, unlabeled network packets, including metadata, payloads, and header information. This data is suitable for training and evaluating machine learning models for network intrusion detection, cybersecurity, and traffic analysis.
HoneyPot repository: hachimi on GitHub.
Dataset Overview
The Internet Background Noise… See the full description on the dataset page: https://huggingface.co/datasets/burpheart/Internet-background-noise.ASR-WPM-And-Background-Noise-Eval
ASR WPM and Background Noise Evaluation Dataset
A dataset of annotated audio recordings for evaluating how different factors affect Whisper (and other ASR/STT systems) transcription accuracy.
Purpose
This dataset provides controlled audio samples with annotations to evaluate ASR performance across:
Speaking pace (fast, normal, slow, mumbled, whispered, weird voices)
Background noise (cafe, music, conversations in various languages, traffic, sirens, etc.)
Microphone… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/ASR-WPM-And-Background-Noise-Eval.background-noise-detection-dataset
Speech-Free Background Noise Dataset — Real-World, Non-Synthetic (50+ Hours)
Dataset summary
50+ hours of real-world urban environmental/ambient background noise (field recordings) without intelligible speech (speech-free), from three scenes: airport, street, subway. The dataset is non-synthetic and intended for speech enhancement via noise augmentation and sound event detection (SED) as “clean background”/negative class
Full version of dataset is… See the full description on the dataset page: https://huggingface.co/datasets/kakadong2018/background-noise-detection-dataset.instruction-background-noise-data-synthetic
Dataset Card for instruction-background-noise-data-synthetic
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/jan-hq/instruction-background-noise-data-synthetic/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/Menlo/instruction-background-noise-data-synthetic.background-noise-detection-dataset
Speech-Free Background Noise Dataset — Real-World, Non-Synthetic (50+ Hours)
Dataset summary
50+ hours of real-world urban environmental/ambient background noise (field recordings) without intelligible speech (speech-free), from three scenes: airport, street, subway. The dataset is non-synthetic and intended for speech enhancement via noise augmentation and sound event detection (SED) as “clean background”/negative class
Full version of dataset is availible… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/background-noise-detection-dataset.
