CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abnajlae /darija-asr-corpus Darija ASR Corpus (dataset-core) Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). This repo contains four source subsets: DODa, DVoice, Wiki, and YouTube. Each subset carries its own upstream license/terms -- see below -- because they are drawn from four different original projects. Subsets Config Rows Audio bundled? Upstream license Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.audioautomatic-speech-recognition10K<n<100K0 likes539 downloads16d agoHugging Face02Williamsanderson /MedQA-Darija-MultiLingual MedQA-Darija-MultiLingual The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija. A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region. Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.audioquestion-answering100K<n<1M4 likes476 downloads5mo agoHugging Face03Jip7e /doda-darija-cosyvoice2 Dataset Card for DODa Moroccan Darija (CosyVoice2 Ready-to-Train) Dataset Summary DODa Moroccan Darija (CosyVoice2 Edition) is a curated, standardized, and tokenized speech dataset engineered specifically for fine-tuning CosyVoice2 on Moroccan Arabic (Darija). While raw audio datasets typically require extensive preprocessing (sample rate normalization, voice activity detection, multi-speaker segmentation, semantic tokenization, speaker embedding extraction, and… See the full description on the dataset page: https://huggingface.co/datasets/Jip7e/doda-darija-cosyvoice2.audiotext-to-speech10K<n<100K1 likes351 downloads26d agoHugging Face04Datasmartly /Darija3-denoisedaudio1K<n<10K0 likes305 downloads1y agoHugging Face05ohsn /darija_yt_2026 darija_yt_2026 Partition upload generated automatically. Namespace: ohsn Repo: ohsn/darija_yt_2026 Video count: 3511 Duration hours: 1565.31 This dataset contains raw audio files and per-video metadata generated from the Darija YouTube extraction pipeline. audioautomatic-speech-recognition1K<n<10K0 likes271 downloads21d agoHugging Face06adiren7 /darija_speech_to_textaudioautomatic-speech-recognition10K<n<100K13 likes180 downloads2y agoHugging Face07ToumAIAnalytics /darija_sttgatedaudio100K<n<1M0 likes156 downloads2mo agoHugging Face08ayoubkirouane /darija-stt-mixThe Darija Speech To Text Dataset is a comprehensive collection designed to support speech recognition tasks for the Darija dialect, it includes audio data totaling 8.23 GB and consists of 13,178 rows of transcribed speech. This dataset covers a variety of dialects, primarily focusing on Algerian and Moroccan Darija, and also includes slang from other Arabic-speaking countries. The data has been meticulously gathered from diverse resources to ensure a rich and varied representation of spoken… See the full description on the dataset page: https://huggingface.co/datasets/ayoubkirouane/darija-stt-mix.audioautomatic-speech-recognition10K<n<100K6 likes154 downloads2y agoHugging Face09ntariklk /darija-merged-asraudio10K<n<100K0 likes145 downloads4mo agoHugging Face1001Yassine /darija-asr-3h Moroccan Darija ASR — 3 hours YouTube Moroccan Darija, segmented and filtered, labeled with Gemini 2.5 Pro. split hours clips train 3.00 1778 validation 0.15 91 silver 0.35 184 Splits are channel-disjoint: silver channels do not appear in train. A same-size random split leaks 100% of silver channels into train. Columns id, audio (16 kHz), text (Gemini 2.5 Pro) channel (YouTube handle) duration, pesq_hyp (SQUIM, no-reference), num_speakers… See the full description on the dataset page: https://huggingface.co/datasets/01Yassine/darija-asr-3h.audioautomatic-speech-recognition1K<n<10K1 likes141 downloads25d agoHugging Face11KandirResearch /DarijaTTS-cleanaudio10K<n<100K1 likes93 downloads2y agoHugging Face12anassdabaghi /tts_darija language: - ar license: cc-by-4.0 task_categories: - automatic-speech-recognition task_ids: - automatic-speech-recognition pretty_name: Darija Arabic Speech Dataset size_categories: - 1K<n<10K tags: - darija - moroccan-arabic - arabic - speech - asr - automatic-speech-recognition - whisper - morocco Moroccan Darija Speech Dataset A speech dataset for Moroccan Arabic (Darija) automatic speech recognition (ASR). The dataset consists of short audio clips extracted… See the full description on the dataset page: https://huggingface.co/datasets/anassdabaghi/tts_darija.audiotext-to-speech10K<n<100K0 likes91 downloads6d agoHugging Face13BrunoHays /DVOICEv2.0-DarijaDVoice is a community initiative that aims to provide African languages and dialects with data and models to facilitate their use of voice technologies. The lack of data on these languages makes it necessary to collect data using methods that are specific to each language. Two different approaches are currently used: the DVoice platform, which is based on Mozilla Common Voice, for collecting authentic recordings from the community, and transfer learning techniques for automatically labeling… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/DVOICEv2.0-Darija.audio10K<n<100K1 likes87 downloads2y agoHugging Face14ai-ssam /darija-tts-8400 Darija TTS 8400 Synthetic Moroccan Darija speech for TTS fine-tuning: 8,400 single-speaker clips (20.73 hours), 24 kHz mono PCM16 WAV. All audio is generated with Gemini 3.1 Flash TTS (gemini-3.1-flash-tts-preview, voice Kore). Clips are unreviewed; there are no human recordings. Write-up of how this data was used: Training a Voice. At a glance Clips / hours 8,400 / 20.73 Unique texts 4,800 Voice Kore (1 speaker) Sample rate 24 kHz mono PCM16… See the full description on the dataset page: https://huggingface.co/datasets/ai-ssam/darija-tts-8400.audiotext-to-speech1K<n<10K0 likes85 downloads8d agoHugging Face15BrunoHays /darija-speech-to-text Speech To Text Darija dataset Reupload of adiren7/darija_speech_to_text audioautomatic-speech-recognition1K<n<10K6 likes83 downloads2y agoHugging Face16abnajlae /darija-asr-benchmark-6speaker Darija ASR 6-Speaker Benchmark A fixed, paired 20-utterance benchmark read identically by 6 held-out speakers (3 female: F1, F2, F3; 3 male: M1, M2, M3 -- none present in any training corpus), used to evaluate cross-speaker generalization for a Moroccan Darija (Arabizi) Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). Consent and anonymization Written informed consent was obtained from all six speakers for the recording and… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-benchmark-6speaker.audioautomatic-speech-recognitionn<1K0 likes68 downloads16d agoHugging Face17hamzasibous /darija-stt-dataset Dataset Card for "darija-stt-dataset" More Information needed audio1K<n<10K0 likes59 downloads1y agoHugging Face18BrunoHays /darija_stt_mixaudio10K<n<100K0 likes50 downloads2y agoHugging Face19adiren7 /darija_to_french_speech_to_textaudion<1K7 likes47 downloads2y agoHugging Face20atlasia /Moroccan-Darija-Wiki-Audio-Datasetgated Moroccan Darija Wiki Audio Dataset Overview The Moroccan Darija Wiki Audio Dataset consists of 551 parallel text and speech samples of Moroccan Darija sourced from Wikipedia Darija . This dataset is designed to support speech recognition, language modeling, and various NLP tasks for Moroccan Darija. Dataset Source The data was scraped from Wikipedia (ary) using the WikiScraper tool. Data Preprocessing To ensure data quality, we applied… See the full description on the dataset page: https://huggingface.co/datasets/atlasia/Moroccan-Darija-Wiki-Audio-Dataset.audion<1K16 likes46 downloads2y agoHugging Face21mohamedmou /DATASET-darija Darija ASR Dataset Dataset de reconnaissance automatique de la parole en Darija marocaine. Description Langue: Darija marocaine (ary) Tache: automatic speech recognition Audio: WAV mono 16 kHz stocke en Parquet Colonnes: audio, sentence Structure Colonne Type Description audio Audio Segment audio WAV mono 16 kHz sentence string Transcription en darija License CC BY 4.0 audioautomatic-speech-recognitionn<1K1 likes46 downloads5mo agoHugging Face22afyfbadreddine77 /darija-asr-cleanaudio10K<n<100K0 likes44 downloads5mo agoHugging Face23fares5462 /darija-dz-tts-v1audion<1K0 likes42 downloads9d agoHugging Face24mohamedmou /moroccan-darija-asr-dataset-splitaudio10K<n<100K1 likes41 downloads5mo agoHugging Face25anaszil /Segmented-Moroccan-Darija-Wiki-Audio-Dataset Dataset Card for Segmented Moroccan Darija Wiki Dataset Dataset Summary This dataset provides short Moroccan Darija (Moroccan Arabic) speech segments derived from the atlasia/Moroccan-Darija-Wiki-Audio-Dataset.It is a cleaned and segmented version of the parent dataset, text-cleaned with Gemini 2.5-flash and processed using a fine-tuned Whisper model for Darija (to be open-sourced soon). Each audio is split into segments of up to 30 seconds to make it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/anaszil/Segmented-Moroccan-Darija-Wiki-Audio-Dataset.audioautomatic-speech-recognition1K<n<10K1 likes39 downloads1y agoHugging Face26igitsml /darija-synthetic-callsaudio1K<n<10K0 likes35 downloads9mo agoHugging Face27mohamedmou /DATASET-darija-ASR-clean Darija ASR Dataset Dataset de reconnaissance automatique de la parole en Darija marocaine. Description Langue: Darija marocaine (ary) Tache: automatic speech recognition Audio: WAV mono 16 kHz stocke en Parquet Colonnes: audio, sentence Structure Colonne Type Description audio Audio Segment audio WAV mono 16 kHz sentence string Transcription en darija License CC BY 4.0 audioautomatic-speech-recognitionn<1K0 likes30 downloads5mo agoHugging Face28BrunoHays /DVOICEv1.1-DarijaDialectal Voice is a community project initiated by AIOX Labs to facilitate voice recognition by Intelligent Systems. Today, the need for AI systems capable of recognizing the human voice is increasingly expressed within communities. However, we note that for some languages such as Darija, there are not enough voice technology solutions. To meet this need, we then proposed to establish this program of iterative and interactive construction of a dialectal database open to all in order to help… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/DVOICEv1.1-Darija.audio1K<n<10K0 likes28 downloads2y agoHugging Face29Fahd1199 /darija-tts Moroccan Darija TTS Dataset This dataset contains Moroccan Darija speech recordings and their corresponding transcriptions, intended for fine-tuning text-to-speech models. Dataset Structure wavs_16k/: Directory containing 16kHz mono 16-bit WAV audio files. metadata_train.csv: CSV file with training data. metadata_val.csv: CSV file with validation data. Usage Load the dataset using: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Fahd1199/darija-tts.audio1K<n<10K1 likes27 downloads1y agoHugging Face30DrIAmed /darija-HF-datasetaudio1K<n<10K0 likes23 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.