CoolFace
Datasetpublicgated

hbredin/pyannoteAI-EMMA-20k

pyannoteAI-EMMA-20k dataset In the framework of the EMMA project, part of JSALT 2025, the pyannoteAI research lab ran its internal speaker diarization pipeline on the whole YODAS2 dataset. From the output, we then selected a subset of the predictions for a total amount of around 20k hours of audio. Licence This dataset is licensed under CC BY-NC-SA 4.0 and therefore does not allow commercial use (e.g. training a commercial model using this dataset).… See the full description on the dataset page: https://huggingface.co/datasets/hbredin/pyannoteAI-EMMA-20k.

sourceHugging Facecc-by-nc-sa-4.0updated 1y agoView on Hugging Face
1likes5downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.