CoolFace
Datasetpublicgated

atlasia/Moroccan-Darija-Wiki-Audio-Dataset

Moroccan Darija Wiki Audio Dataset Overview The Moroccan Darija Wiki Audio Dataset consists of 551 parallel text and speech samples of Moroccan Darija sourced from Wikipedia Darija . This dataset is designed to support speech recognition, language modeling, and various NLP tasks for Moroccan Darija. Dataset Source The data was scraped from Wikipedia (ary) using the WikiScraper tool. Data Preprocessing To ensure data quality… See the full description on the dataset page: https://huggingface.co/datasets/atlasia/Moroccan-Darija-Wiki-Audio-Dataset.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
16likes46downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.