datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1
serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1
This is the first pushed Ghanaian Speech Lab ASR pipeline artifact. It is a
review artifact for the v0.1 Akan ASR pass, not a trained model checkpoint.
Expected future model repo:
teckedd/serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1
What This Artifact Contains
data/manifest.jsonl: harmonized Waxal + GhanaNLP manifest references.
reports/sanitize.json: sanitization report and… See the full description on the dataset page: https://huggingface.co/datasets/teckedd/serendepify-gsl-asr-ak-waxal-gnlp-whisper-small-replay-fullft-v0.1.arabic-whisper-multidialect-processed-small
Arabic Whisper Multi-Dialect - Processed (Small)
Dataset Description
This is a preprocessed version of the Arabic multi-dialect speech dataset, ready for fine-tuning OpenAI's Whisper models. The dataset contains audio features extracted and formatted specifically for Whisper training.
Size: 40% subset of the full arabic-whisper-multidialect dataset
Total Examples: 43,091 samples
Format: Pre-computed Whisper input features (mel spectrograms) and tokenized labels
Purpose:… See the full description on the dataset page: https://huggingface.co/datasets/MadLook/arabic-whisper-multidialect-processed-small.
