hbredin/pyannoteAI-EMMA-20k
pyannoteAI-EMMA-20k dataset In the framework of the EMMA project, part of JSALT 2025, the pyannoteAI research lab ran its internal speaker diarization pipeline on the whole YODAS2 dataset. From the output, we then selected a subset of the predictions for a total amount of around 20k hours of audio. Licence This dataset is licensed under CC BY-NC-SA 4.0 and therefore does not allow commercial use (e.g. training a commercial model using this dataset).… See the full description on the dataset page: https://huggingface.co/datasets/hbredin/pyannoteAI-EMMA-20k.
This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.
