datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yodas-en-replay
YODAS-EN Replay
General-domain English speech for replay mixing during domain adaptation, with
cased and punctuated transcripts. Audio comes from
espnet/yodas2 (CC-BY-3.0, sourced
from Creative-Commons YouTube videos); the transcripts are our own, produced with
faster-whisper large-v3-turbo.
Why this exists
If you fine-tune a small ASR model on a narrow domain, it forgets everything else.
We measured 5 hours of meeting audio buying 0.74 WER points in-domain while… See the full description on the dataset page: https://huggingface.co/datasets/moonshine-ai/yodas-en-replay.serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1
serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1
This is the first pushed Ghanaian Speech Lab ASR pipeline artifact. It is a
review artifact for the v0.1 Akan ASR pass, not a trained model checkpoint.
Expected future model repo:
teckedd/serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1
What This Artifact Contains
data/manifest.jsonl: harmonized Waxal + GhanaNLP manifest references.
reports/sanitize.json: sanitization report and… See the full description on the dataset page: https://huggingface.co/datasets/teckedd/serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1.
