atlasia/Moroccan-Darija-Wiki-Audio-Dataset
Moroccan Darija Wiki Audio Dataset Overview The Moroccan Darija Wiki Audio Dataset consists of 551 parallel text and speech samples of Moroccan Darija sourced from Wikipedia Darija . This dataset is designed to support speech recognition, language modeling, and various NLP tasks for Moroccan Darija. Dataset Source The data was scraped from Wikipedia (ary) using the WikiScraper tool. Data Preprocessing To ensure data quality… See the full description on the dataset page: https://huggingface.co/datasets/atlasia/Moroccan-Darija-Wiki-Audio-Dataset.
1646
No card is published for this repository, or it could not be fetched from Hugging Face right now.
