datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
atcosim-speaker-disjoint-splits
ATCOSIM speaker-disjoint splits (metadata only)
This dataset contains no audio and no transcripts. It is a split definition:
one row per ATCOSIM utterance, giving its speaker, its recording session, its
duration, and which half of a speaker-disjoint evaluation it belongs to.
The audio and transcriptions are not here because they cannot be redistributed.
The ATCOSIM corpus
manual §5.2 states that the corpus is "provided free of charge" and "permitted
to use ... for research and… See the full description on the dataset page: https://huggingface.co/datasets/moonshine-ai/atcosim-speaker-disjoint-splits.uwb-atcc-session-disjoint-splits
UWB-ATCC session-disjoint splits
IDs only, no audio. Train drops the one session shared with the published test split (TWR-34720N) so an in-domain number is domain adaptation, not session leakage. Audio stays on Jzuluaga/uwb_atcc (CC BY-NC-SA 4.0).
