bleep
Datasets
All datasets matching “bleep”bleep-spans
Bleep spans — synthetic sensitive-speech regions with frame-accurate labels
Where sensitive information is spoken, and what kind it is — never what was
said.
Every recording is synthetic. No real telephone call, clinical recording, or any
other real speech was used, recorded, or derived from at any stage.
🤗 Model: NagaYu/bleep-0.09b
🎛️ Demo: NagaYu/bleep
What a row contains
utt_id, voice_key, condition, duration, subsets, and three parallel
arrays —… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/bleep-spans.bleep
BLEEP — Broadcast Language Elicitation and Evaluation for Profanity
v2.0 · 8,312 clips · 86 speakers · English (US/UK) · 16 kHz mono · 4.62 h
Isolated-word English corpus for profanity speech research. 20 profanity keywords and 29 hard
negatives — minimal-pair confusables selected by CMUdict phoneme edit distance and SUBTLEX-US/UK
frequency. Each speaker recorded all 49 words twice in one session, once in a neutral and
once in an expressive register.
⚠️ Content warning — every… See the full description on the dataset page: https://huggingface.co/datasets/ritinrk/bleep.reddit
