datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Common-Voice-17-Arabic-for-Seasme-CSM-Finetuning
Curated Arabic Speech Dataset for Seasme (from MCV17)
Dataset Description
This dataset is a curated and preprocessed version of the Arabic (ar) subset from Mozilla Common Voice (MCV) 17.0. It has been specifically prepared for fine-tuning conversational speech models, with a primary focus on the Seasme-CSM model architecture. The dataset consists of audio clips in WAV format (24kHz, mono) and their corresponding transcripts, along with integer speaker IDs.
The original… See the full description on the dataset page: https://huggingface.co/datasets/MAdel121/Common-Voice-17-Arabic-for-Seasme-CSM-Finetuning.how-people-make-money-csm1bfinetuned-lb-ar-csm-3-5h-groupedcsm-turkish-ttsfinetuned-lb-ar-csm-3-1h-groupedmorgan-csmfinetuned-lb-ar-csm-3-2h-groupedfinetuned-lb-ar-csm-3-full-groupedza-zenande-csm-metadata2-v1us-julia-csm-v1finetuned-lb-ar-csm-3-3h-groupedfinetuned-lb-ar-csm-3-7h-grouped
