CoolFace
Datasetpublic

mexus/ru-book-mix-10h

ru-book-mix-10h A 10-hour synthetic Russian-audiobook diarization benchmark. 600 one-minute FLAC clips (16 kHz mono, 16-bit, lossless) with NIST RTTM ground truth, generated by mexus/diarization-benchmark from its5Q/biggest-ru-book (speech) and bilguun/musan-noise (background). Intended use: diarization evaluation only. This dataset is not suitable for training — the same source voices repeat across files, so any model that trains on it will leak voice identity into its test… See the full description on the dataset page: https://huggingface.co/datasets/mexus/ru-book-mix-10h.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes108downloads
2 commits on main
18655c44mo ago

Add 10h diarization benchmark: 4 FLAC WebDataset shards + manifest + README

mexus
3e33a634mo ago

initial commit

mexus