datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CosyVoice2-SparkTTSdoda-darija-cosyvoice2
Dataset Card for DODa Moroccan Darija (CosyVoice2 Ready-to-Train)
Dataset Summary
DODa Moroccan Darija (CosyVoice2 Edition) is a curated, standardized, and tokenized speech dataset engineered specifically for fine-tuning CosyVoice2 on Moroccan Arabic (Darija).
While raw audio datasets typically require extensive preprocessing (sample rate normalization, voice activity detection, multi-speaker segmentation, semantic tokenization, speaker embedding extraction, and… See the full description on the dataset page: https://huggingface.co/datasets/Jip7e/doda-darija-cosyvoice2.cosyvoice2_enCosyVoice2synth-cosyvoice2
