CoolFace
Datasetpublic

eko57/spc_r_segmented

eko57/spc_r_segmented Diarized and segmented speech dataset derived from i4ds/spc_r. Description Each row is a merged speech segment belonging to a single speaker. The source audio and SRT subtitles from i4ds/spc_r were processed with the following pipeline: Diarization -- pyannote/speaker-diarization-3.1 assigned speaker labels to each SRT segment based on temporal overlap. Merging -- Consecutive SRT segments from the same speaker were merged when the silence… See the full description on the dataset page: https://huggingface.co/datasets/eko57/spc_r_segmented.

sourceHugging Facecc-by-4.0updated 7mo agoView on Hugging Face
0likes89downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
eko57/spc_r_segmented · CoolFace