noise
noisejan-hq_-_llama3-s-instruct-v0.3-checkpoint-7000-phase-3-10k-noise-ggufTakvmi_-_model_pmc_gamma_0.05_noise0.2_0.1_epoch0-ggufTakvmi_-_model_pmc_gamma_0.05_noise0.2_0.2_epoch2-ggufTakvmi_-_model_pmc_gamma_0.05_noise0.1_0.1_epoch0-ggufTakvmi_-_model_pmc_gamma_0.05_noise0.2_0.1_epoch1-ggufwan2.2_i2v_A14b_high_noise_lora_rank64_lightx2v_4step_1022Takvmi_-_model_pmc_gamma_0.1_noise0.2_0.2_epoch0-gguf
Datasets
All datasets matching “noise”Vaani-Noise-Event-Dataset
Vaani Noise Event Timestamps
Dataset Summary
Vaani Noise Event Timestamps is a derived dataset from Project Vaani, a large-scale multilingual speech initiative by IISc Bangalore and ARTPARK that captures India's linguistic diversity across all districts.
This dataset provides noise event annotations with fine-grained timestamps for the subset audio recordings from the Vaani corpus. Each entry identifies background noise categories along with their precise start… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-Noise-Event-Dataset.librispeech_asr-noise
Dataset Card for "librispeech_asr-noise"
More Information needed
ImageNet-C-gaussian_noise-severity_5generative-sound-masking-input-noise-full-v1
Generative Sound Masking input-noise pool v1
This WebDataset contains 48,840 mono 16-kHz, 10.24-second input-noise clips
across 49 tar shards. It combines the complete Yiming SONYC, TAU Urban
Acoustic Scenes, and UrbanSound baseline with subject-balanced BABYCRY-UJM-AXA
and NOTSOFAR-1 train windows. Stable sample metadata are in
metadata/noise_index.jsonl; JSON beside each WAV adds hashes computed during
packaging.
The source datasets carry different licenses. In particular… See the full description on the dataset page: https://huggingface.co/datasets/AE-W/generative-sound-masking-input-noise-full-v1.ambient_noise_audio
Ambient Noise Collection dataset
This dataset is useful for reducing ASR model Hallucinations (especially Whisper), by default Whisper often transcribing non-speech audio as hallucination transcriptions.
This dataset attempts to improve ASR model that have hallucinations on non-speech audio.
Collected several audio from source:
https://www.kaggle.com/datasets/nafin59/hospital-ambient-noise
https://www.kaggle.com/datasets/solorzano/ambient-noise… See the full description on the dataset page: https://huggingface.co/datasets/Willy030125/ambient_noise_audio.OLMo-2-2.7B-Exp-NoiseVectors
OLMo-2-2.7B-Exp Noise Vectors
Gaussian noise vectors added to the input embeddings during pretraining of
sbordt/OLMo-2-2.7B-Exp
(a 2.7B-parameter OLMo-2-style model with d_model=2880). Released as a
uniform-random 1% subsample per every-1000-batch chunk from 51,200 poisoned
pretraining batches over 100,000 training steps — 480 rows total.
How the noise was applied during training
For each poisoned batch, Gaussian noise of shape (4096, 2880) was
drawn and added to… See the full description on the dataset page: https://huggingface.co/datasets/sbordt/OLMo-2-2.7B-Exp-NoiseVectors.
