CoolFace
Datasetpublic

vnahata/IndicDiarBench-speaker-retrieval

Indic DiarBench speaker retrieval (MTEB) Indic DiarBench reshaped for speaker retrieval across the 22 scheduled languages of India: given a clip of one speaker, find other clips of that same speaker. Source: sarvamai/indic-diarbench at revision 92877ba, cc-by-4.0, official test split. Turns are cut by their annotated times, restricted to 2 to 15 seconds, and turns overlapping a different speaker are dropped. identity pairs the recording session with the speaker, because speaker… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/IndicDiarBench-speaker-retrieval.

sourceHugging Facecc-by-4.0updated 22d agoView on Hugging Face
0likes509downloads
Dataset Card

Indic DiarBench speaker retrieval (MTEB)

Indic DiarBench reshaped for speaker retrieval across the 22 scheduled languages of India: given a clip of one speaker, find other clips of that same speaker.

Source: sarvamai/indic-diarbench at revision 92877ba, cc-by-4.0, official test split. Turns are cut by their annotated times, restricted to 2 to 15 seconds, and turns overlapping a different speaker are dropped. identity pairs the recording session with the speaker, because speaker labels are numbered per session.

Built by scripts/data/indic_diarbench_speaker/create_data.py in the MTEB repo.