CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Myrtle /CAIMAN-ASR-BackgroundNoise Dataset Card for Myrtle/CAIMAN-ASR-BackgroundNoise This dataset provides background noise audio, suitable for noise augmentation while training Myrtle.ai's CAIMAN-ASR models. Dataset Details Dataset Description Curated by: Myrtle.ai License: Myrtle.ai's modifications to the source data are licensed under the CC BY 4.0 license. Some of the original data is under the CC BY 3.0 license; the rest is in the public domain. Please see the Source Data section… See the full description on the dataset page: https://huggingface.co/datasets/Myrtle/CAIMAN-ASR-BackgroundNoise.audio1K<n<10K11 likes1.1k downloads3y agoHugging Face02burpheart /Internet-background-noise Internet Background Noise Dataset (Unlabeled Raw Data) This dataset contains HTTP internet noise data collected by an internet honeypot. It consists of raw, unlabeled network packets, including metadata, payloads, and header information. This data is suitable for training and evaluating machine learning models for network intrusion detection, cybersecurity, and traffic analysis. HoneyPot repository: hachimi on GitHub. Dataset Overview The Internet Background Noise… See the full description on the dataset page: https://huggingface.co/datasets/burpheart/Internet-background-noise.tabulartext-classification1M<n<10M6 likes42 downloads2y agoHugging Face03danielrosehill /ASR-WPM-And-Background-Noise-Eval ASR WPM and Background Noise Evaluation Dataset A dataset of annotated audio recordings for evaluating how different factors affect Whisper (and other ASR/STT systems) transcription accuracy. Purpose This dataset provides controlled audio samples with annotations to evaluate ASR performance across: Speaking pace (fast, normal, slow, mumbled, whispered, weird voices) Background noise (cafe, music, conversations in various languages, traffic, sirens, etc.) Microphone… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/ASR-WPM-And-Background-Noise-Eval.audioautomatic-speech-recognitionn<1K1 likes32 downloads10mo agoHugging Face04kakadong2018 /background-noise-detection-dataset Speech-Free Background Noise Dataset — Real-World, Non-Synthetic (50+ Hours) Dataset summary 50+ hours of real-world urban environmental/ambient background noise (field recordings) without intelligible speech (speech-free), from three scenes: airport, street, subway. The dataset is non-synthetic and intended for speech enhancement via noise augmentation and sound event detection (SED) as “clean background”/negative class Full version of dataset is… See the full description on the dataset page: https://huggingface.co/datasets/kakadong2018/background-noise-detection-dataset.audion<1K1 likes27 downloads2mo agoHugging Face05Menlo /instruction-background-noise-data-synthetic Dataset Card for instruction-background-noise-data-synthetic This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/jan-hq/instruction-background-noise-data-synthetic/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/Menlo/instruction-background-noise-data-synthetic.text10K<n<100K0 likes19 downloads2y agoHugging Face06AxonData /background-noise-detection-dataset Speech-Free Background Noise Dataset — Real-World, Non-Synthetic (50+ Hours) Dataset summary 50+ hours of real-world urban environmental/ambient background noise (field recordings) without intelligible speech (speech-free), from three scenes: airport, street, subway. The dataset is non-synthetic and intended for speech enhancement via noise augmentation and sound event detection (SED) as “clean background”/negative class Full version of dataset is availible… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/background-noise-detection-dataset.audion<1K1 likes17 downloads7mo agoHugging Face07NashAli /background-noise-detection-dataset Speech-Free Background Noise Dataset — Real-World, Non-Synthetic (50+ Hours) Dataset summary 50+ hours of real-world urban environmental/ambient background noise (field recordings) without intelligible speech (speech-free), from three scenes: airport, street, subway. The dataset is non-synthetic and intended for speech enhancement via noise augmentation and sound event detection (SED) as “clean background”/negative class Purpose and usage scenarios Speech… See the full description on the dataset page: https://huggingface.co/datasets/NashAli/background-noise-detection-dataset.audion<1K0 likes15 downloads8mo agoHugging Face08kdcyberdude /background_noiseaudio1 likes12 downloads2y agoHugging Face09Bear3 /background_noiseaudio0 likes10 downloads1y agoHugging Face10jan-hq /instruction-background-noise-datatext1K<n<10K0 likes9 downloads2y agoHugging Face11ClarusC64 /clinical-quad-ae-signal-background-noise-reporting-lag-causality-bias-v0.1Clinical Quad AE Noise Lag Attribution Bias v0.1 Each row is a site week safety snapshot. Core quad AE signal rateBackground noise rateReporting lagAttribution bias Target label_stop_signal_next_30d Files data/train.csvdata/tester.csvscorer.py Evaluation Run model on data/tester.csvReturn predictions row alignedScore with scorer.py License MIT tabulartext-classificationn<1K0 likes6 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.