moonshine
audio_samples_1kyodas-en-replay
YODAS-EN Replay
General-domain English speech for replay mixing during domain adaptation, with
cased and punctuated transcripts. Audio comes from
espnet/yodas2 (CC-BY-3.0, sourced
from Creative-Commons YouTube videos); the transcripts are our own, produced with
faster-whisper large-v3-turbo.
Why this exists
If you fine-tune a small ASR model on a narrow domain, it forgets everything else.
We measured 5 hours of meeting audio buying 0.74 WER points in-domain while… See the full description on the dataset page: https://huggingface.co/datasets/moonshine-ai/yodas-en-replay.multilingual_examplesmoonshine-streamingatcosim-speaker-disjoint-splits
ATCOSIM speaker-disjoint splits (metadata only)
This dataset contains no audio and no transcripts. It is a split definition:
one row per ATCOSIM utterance, giving its speaker, its recording session, its
duration, and which half of a speaker-disjoint evaluation it belongs to.
The audio and transcriptions are not here because they cannot be redistributed.
The ATCOSIM corpus
manual §5.2 states that the corpus is "provided free of charge" and "permitted
to use ... for research and… See the full description on the dataset page: https://huggingface.co/datasets/moonshine-ai/atcosim-speaker-disjoint-splits.uwb-atcc-session-disjoint-splits
UWB-ATCC session-disjoint splits
IDs only, no audio. Train drops the one session shared with the published test split (TWR-34720N) so an in-domain number is domain adaptation, not session leakage. Audio stays on Jzuluaga/uwb_atcc (CC BY-NC-SA 4.0).
