MadLook/arabic-whisper-multidialect-processed-small
Arabic Whisper Multi-Dialect - Processed (Small) Dataset Description This is a preprocessed version of the Arabic multi-dialect speech dataset, ready for fine-tuning OpenAI's Whisper models. The dataset contains audio features extracted and formatted specifically for Whisper training. Size: 40% subset of the full arabic-whisper-multidialect dataset Total Examples: 43,091 samples Format: Pre-computed Whisper input features (mel spectrograms) and tokenized labels… See the full description on the dataset page: https://huggingface.co/datasets/MadLook/arabic-whisper-multidialect-processed-small.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face