CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MBZUAI-Paris /DarijaMMLU Dataset Card for DarijaMMLU Dataset Summary DarijaMMLU is an evaluation benchmark designed to assess large language models' (LLM) performance in Moroccan Darija, a variety of Arabic. It consists of 22,027 multiple-choice questions, translated from selected subsets of the Massive Multitask Language Understanding (MMLU) and ArabicMMLU benchmarks to measure model performance on 44 subjects in Darija. Supported Tasks Task Category: Multiple-choice question… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/DarijaMMLU.textquestion-answering10K<n<100K8 likes3k downloads2y agoHugging Face02touati-kamel /TinyStories-Algerian-Darijatabular10K<n<100K0 likes2.3k downloads18d agoHugging Face03imomayiz /darija-englishThis work is part of DODa. texttranslation10K<n<100K13 likes506 downloads2y agoHugging Face04abnajlae /darija-asr-corpus Darija ASR Corpus (dataset-core) Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). This repo contains four source subsets: DODa, DVoice, Wiki, and YouTube. Each subset carries its own upstream license/terms -- see below -- because they are drawn from four different original projects. Subsets Config Rows Audio bundled? Upstream license Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.audioautomatic-speech-recognition10K<n<100K0 likes473 downloads15d agoHugging Face05Williamsanderson /MedQA-Darija-MultiLingual MedQA-Darija-MultiLingual The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija. A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region. Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.audioquestion-answering100K<n<1M4 likes453 downloads5mo agoHugging Face06nasrellahkharroubi /DarijaDz DarijaDZ DarijaDZ is a large-scale corpus of user-generated text collected from Algerian YouTube and TikTok channels. The corpus contains approximately 22.5 million comment documents and 259.22 million word-level tokens, with content written primarily in Algerian Darija script alongside Latin/Arabizi writing and mixed-script content. Dataset Description Motivation Algerian Darija is an under-resourced language variety with comparatively limited… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDz.texttext-generation10M<n<100M10 likes334 downloads8d agoHugging Face07Jip7e /doda-darija-cosyvoice2 Dataset Card for DODa Moroccan Darija (CosyVoice2 Ready-to-Train) Dataset Summary DODa Moroccan Darija (CosyVoice2 Edition) is a curated, standardized, and tokenized speech dataset engineered specifically for fine-tuning CosyVoice2 on Moroccan Arabic (Darija). While raw audio datasets typically require extensive preprocessing (sample rate normalization, voice activity detection, multi-speaker segmentation, semantic tokenization, speaker embedding extraction, and… See the full description on the dataset page: https://huggingface.co/datasets/Jip7e/doda-darija-cosyvoice2.audiotext-to-speech10K<n<100K1 likes334 downloads24d agoHugging Face08ayoubkirouane /Algerian-Darija Overview This dataset contains text in Algerian Darija, collected from a variety of sources including existing datasets on Hugging Face, web scraping, and YouTube transcript APIs. The train split consists more then 2k rows of uncleaned text data. The v1 split consists more than 170k rows of split and partially cleaned text. Sources The text data was gathered from: Hugging Face Datasets: Pre-existing datasets relevant to Algerian Darija. Web Scraping: Content… See the full description on the dataset page: https://huggingface.co/datasets/ayoubkirouane/Algerian-Darija.texttext-generation100K<n<1M16 likes292 downloads5mo agoHugging Face09Datasmartly /Darija3-denoisedaudio1K<n<10K0 likes292 downloads1y agoHugging Face10ohsn /darija_yt_2026 darija_yt_2026 Partition upload generated automatically. Namespace: ohsn Repo: ohsn/darija_yt_2026 Video count: 3511 Duration hours: 1565.31 This dataset contains raw audio files and per-video metadata generated from the Darija YouTube extraction pipeline. audioautomatic-speech-recognition1K<n<10K0 likes269 downloads20d agoHugging Face11MBZUAI-Paris /DarijaHellaSwag Dataset Card for DarijaHellaSwag Dataset Summary DarijaHellaSwag is a challenging multiple-choice benchmark designed to evaluate machine reading comprehension and commonsense reasoning in Moroccan Darija. It is a translated version of the HellaSwag validation set, which presents scenarios where models must choose the most plausible continuation of a passage from four options. Supported Tasks Task Category: Multiple-choice question answering Task: Answering… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/DarijaHellaSwag.textquestion-answering10K<n<100K4 likes256 downloads2y agoHugging Face12MBZUAI-Paris /DarijaBench DarijaBench: A Comprehensive Evaluation Dataset for Summarization, Translation, and Sentiment Analysis in Darija Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact. The Moroccan Arabic dialect, commonly referred to as Darija, is a widely spoken but understudied variant of Arabic with distinct linguistic features that differ… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/DarijaBench.text10K<n<100K4 likes191 downloads2y agoHugging Face13ToumAIAnalytics /darija_sttgatedaudio100K<n<1M0 likes169 downloads2mo agoHugging Face14ayoubkirouane /darija-stt-mixThe Darija Speech To Text Dataset is a comprehensive collection designed to support speech recognition tasks for the Darija dialect, it includes audio data totaling 8.23 GB and consists of 13,178 rows of transcribed speech. This dataset covers a variety of dialects, primarily focusing on Algerian and Moroccan Darija, and also includes slang from other Arabic-speaking countries. The data has been meticulously gathered from diverse resources to ensure a rich and varied representation of spoken… See the full description on the dataset page: https://huggingface.co/datasets/ayoubkirouane/darija-stt-mix.audioautomatic-speech-recognition10K<n<100K6 likes157 downloads2y agoHugging Face15algerian-nlp /TinyStories-Algerian-Darija TinyStories Algerian Darija Synthetic short stories in Algerian Darja paired with their English originals, for Darija language modeling and translation, from the Algerian NLP Collective. The Hub datasets-server reports 11,326 train rows (/info?dataset=algerian-nlp/TinyStories-Algerian-Darija, 2026-09-17), independently confirmed by the build ledger processed_story_ids.json in this repo: 11,326 unique story ids (0 to 11,492, non-contiguous). The default config answers: what does… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/TinyStories-Algerian-Darija.tabulartext-generation10K<n<100K0 likes154 downloads5d agoHugging Face16adiren7 /darija_speech_to_textaudioautomatic-speech-recognition10K<n<100K13 likes144 downloads2y agoHugging Face17ntariklk /darija-merged-asraudio10K<n<100K0 likes141 downloads4mo agoHugging Face1801Yassine /darija-asr-3h Moroccan Darija ASR — 3 hours YouTube Moroccan Darija, segmented and filtered, labeled with Gemini 2.5 Pro. split hours clips train 3.00 1778 validation 0.15 91 silver 0.35 184 Splits are channel-disjoint: silver channels do not appear in train. A same-size random split leaks 100% of silver channels into train. Columns id, audio (16 kHz), text (Gemini 2.5 Pro) channel (YouTube handle) duration, pesq_hyp (SQUIM, no-reference), num_speakers… See the full description on the dataset page: https://huggingface.co/datasets/01Yassine/darija-asr-3h.audioautomatic-speech-recognition1K<n<10K1 likes139 downloads23d agoHugging Face19bourbouh /moroccan-darija-youtube-subtitles Moroccan Darija YouTube Subtitles Dataset This dataset contains subtitles from YouTube videos in Moroccan Darija, a colloquial Arabic dialect spoken in Morocco. The subtitles were collected from several popular Moroccan YouTube channels, providing a diverse set of transcriptions in the Darija language. Dataset Description The dataset is provided as a CSV file, where each row represents a YouTube video and contains the following columns: video_id: The unique identifier of… See the full description on the dataset page: https://huggingface.co/datasets/bourbouh/moroccan-darija-youtube-subtitles.textothern<1K3 likes115 downloads2y agoHugging Face20nasrellahkharroubi /DarijaDZ-DialectID DarijaDZ Dialect Identification DarijaDZ-DialectID is a labeled dataset for classifying Algerian online text into one of six dialect/language classes: darija, msa, arabize, french, english, code_switch. It is part of DarijaDZ, an attempt to build an NLP ecosystem for Algerian Darija. Dataset Description Motivation Algeria's online text is a mix of several dialects and scripts -- Algerian Darija (Arabic script), Modern Standard Arabic, Arabizi… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDZ-DialectID.texttext-classification10K<n<100K0 likes106 downloads12d agoHugging Face21atlasia /darija_english Dataset Card for atlasia/darija-english Dataset Details Dataset Description A compilation of Darija-English pairs curated by AtlasIA. Curated by: AtlasIA Language(s) (NLP): Moroccan Darija, English License: CC-by-NC-4.0 Darija sentences sources (additionally to the web): doda: AtlasIA platform contributions stories: Mixed Arabic Datasets transliteration: AtlasIA x DODa. Can be used for transliteration task. tabulartranslation100K<n<1M14 likes105 downloads2y agoHugging Face22BrunoHays /DVOICEv2.0-DarijaDVoice is a community initiative that aims to provide African languages and dialects with data and models to facilitate their use of voice technologies. The lack of data on these languages makes it necessary to collect data using methods that are specific to each language. Two different approaches are currently used: the DVoice platform, which is based on Mozilla Common Voice, for collecting authentic recordings from the community, and transfer learning techniques for automatically labeling… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/DVOICEv2.0-Darija.audio10K<n<100K1 likes94 downloads2y agoHugging Face23Lyte /DarijaTTS-v0.2 How to Use the DarijaTTS-v0.2 Dataset Code: import IPython.display as ipd import io import numpy as np import tempfile import wave import os from datasets import load_dataset from IPython.display import Audio # Load the DarijaTTS-v0.2 dataset with streaming streaming_dataset = load_dataset("Lyte/DarijaTTS-v0.2", streaming=True) print("Dataset loaded with streaming:") print(streaming_dataset) # Function to play audio from a streaming dataset element def… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/DarijaTTS-v0.2.text10K<n<100K1 likes93 downloads11mo agoHugging Face24KandirResearch /DarijaTTS-cleanaudio10K<n<100K1 likes92 downloads2y agoHugging Face25atlasia /english-to-darija-arabic-script-formattedtext100K<n<1M0 likes92 downloads3mo agoHugging Face26RanaGaber /FineTranslations_Darijatabular100K<n<1M0 likes86 downloads1mo agoHugging Face27ai-ssam /darija-tts-8400 Darija TTS 8400 Synthetic Moroccan Darija speech for TTS fine-tuning: 8,400 single-speaker clips (20.73 hours), 24 kHz mono PCM16 WAV. All audio is generated with Gemini 3.1 Flash TTS (gemini-3.1-flash-tts-preview, voice Kore). Clips are unreviewed; there are no human recordings. Write-up of how this data was used: Training a Voice. At a glance Clips / hours 8,400 / 20.73 Unique texts 4,800 Voice Kore (1 speaker) Sample rate 24 kHz mono PCM16… See the full description on the dataset page: https://huggingface.co/datasets/ai-ssam/darija-tts-8400.audiotext-to-speech1K<n<10K0 likes85 downloads6d agoHugging Face28anassdabaghi /tts_darija language: - ar license: cc-by-4.0 task_categories: - automatic-speech-recognition task_ids: - automatic-speech-recognition pretty_name: Darija Arabic Speech Dataset size_categories: - 1K<n<10K tags: - darija - moroccan-arabic - arabic - speech - asr - automatic-speech-recognition - whisper - morocco Moroccan Darija Speech Dataset A speech dataset for Moroccan Arabic (Darija) automatic speech recognition (ASR). The dataset consists of short audio clips extracted… See the full description on the dataset page: https://huggingface.co/datasets/anassdabaghi/tts_darija.audiotext-to-speech10K<n<100K0 likes82 downloads4d agoHugging Face29BrunoHays /darija-speech-to-text Speech To Text Darija dataset Reupload of adiren7/darija_speech_to_text audioautomatic-speech-recognition1K<n<10K6 likes80 downloads2y agoHugging Face30abdeljalilELmajjodi /darija_pairs_multilang_datasettext1M<n<10M0 likes78 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.