CoolFace
Datasetpublic

AE-W/generative-sound-masking-generated-energy-v7

Generative Sound Masking — fixed-background audio masking Incrementally generated unfiltered candidates. This is not a final selected dataset. Each background has 15 separately generated prompt–seed outputs using gain-compensated reconstruction residuals. run_config.json pins models, source pools, parameters and implementation hashes. For multiple workers read workers/worker-NN/progress.json; each worker reports only its assigned IDs. Global completion requires all worker… See the full description on the dataset page: https://huggingface.co/datasets/AE-W/generative-sound-masking-generated-energy-v7.

sourceHugging Faceupdated 8m agoView on Hugging Face
0likes5.1kdownloads
Dataset Card

Generative Sound Masking — fixed-background audio masking

Incrementally generated unfiltered candidates. This is not a final selected dataset. Each background has 15 separately generated prompt–seed outputs using gain-compensated reconstruction residuals. runconfig.json pins models, source pools, parameters and implementation hashes. For multiple workers read workers/worker-NN/progress.json; each worker reports only its assigned IDs. Global completion requires all worker reports to be complete. Assignments are backgroundid modulo numworkers. Per-worker runconfig.json pins the parameters. Workers never share a background ID.

Each data/group-NNN/noise-NNNNNN.tar contains one normalized background.wav, 15 rank-NN.mask.wav / rank-NN.mixture.wav pairs and per-candidate JSON, plus manifest.json. Audio is float32 WAV without clipping. Foreground-only gain search targets the configured mixture LUFS. If a BS.1770 gate discontinuity leaves no gain inside tolerance, the nearest measured value is stored and explicitly flagged in candidate metadata. Mixture = calibrated foreground + fixed background at its configured LUFS target. Per-step and final LUFS are recorded. Manifest hashes refer to the exact WAV files; tar hashes and counts are in manifests/. Completed backgrounds are uploaded in groups of 10 to respect Hub commit limits. Confirmed archives are then deleted locally. The final partial group is also uploaded. The sampling step count is pinned in each worker configuration and archive manifest. An interrupted background resumes from its already saved candidates. No PE-AV/CLAP/BEATs/PaSST or audio-LLM validation has been performed.

AE-W/generative-sound-masking-generated-energy-v7 · CoolFace