CoolFace
Datasetpublic

DimitarV/eurospeech-bg-diar-mixtures

⚠️ DEPRECATED — use v2 This dataset contains a shortcut that lets a model infer the number of speakers without listening to the audio. Each speaker was given a fixed 4 turns, so session duration is a direct function of speaker count. Measured on this data: 1-spk median 18.2 s range 11.1-25.2 2-spk 30.4 s 23.0-36.6 3-spk 41.6 s 21.0-54.5 4-spk 54.6 s 38.0-73.0 The 2-speaker and 4-speaker ranges do not overlap — 2-spk tops out… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/eurospeech-bg-diar-mixtures.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
0likes359downloads
Dataset Card
# ⚠️ DEPRECATED — use v2 This dataset contains a shortcut that lets a model infer the number of speakers without listening to the audio. Each speaker was given a fixed 4 turns, so session duration is a direct function of speaker count. Measured on this data: `` 1-spk median 18.2 s range 11.1-25.2 2-spk 30.4 s 23.0-36.6 3-spk 41.6 s 21.0-54.5 4-spk 54.6 s 38.0-73.0 `` The 2-speaker and 4-speaker ranges do not overlap — 2-spk tops out at 36.6 s, 4-spk starts at 40.3 s. Gradient descent will take that free signal, and it does not transfer: a real conversation's length says nothing about how many people are in the room. v2 budgets turns per SESSION rather than per speaker, cutting the spread of medians from 36.4 s to 9.8 s with fully overlapping ranges. It also raises cameo coverage from 11% to 16.4%. Kept for reproducibility and comparison. Do not train on it.