datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
darija-asr-corpus
Darija ASR Corpus (dataset-core)
Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a
Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming).
This repo contains four source subsets: DODa, DVoice, Wiki, and
YouTube. Each subset carries its own upstream license/terms -- see below --
because they are drawn from four different original projects.
Subsets
Config
Rows
Audio bundled?
Upstream license
Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.TORGO-database
The TORGO Database: Acoustic and articulatory speech from speakers with dysarthria
Dataset Summary
This database only includes the short words and restricted sentence portion of the TORGO dataset.
For the full dataset which also includes non-words and unrestricted sentences please see: https://www.cs.toronto.edu/~complingweb/data/TORGO/torgo.html.
Transcripts have been normalized to remove punctuation but casing has been left. Few transcripts only had 'xxx' as text… See the full description on the dataset page: https://huggingface.co/datasets/abnerh/TORGO-database.darija-asr-benchmark-6speaker
Darija ASR 6-Speaker Benchmark
A fixed, paired 20-utterance benchmark read identically by 6 held-out speakers (3
female: F1, F2, F3; 3 male: M1, M2, M3 -- none present in any training corpus),
used to evaluate cross-speaker generalization for a Moroccan Darija (Arabizi)
Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming).
Consent and anonymization
Written informed consent was obtained from all six speakers for the recording and… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-benchmark-6speaker.
