CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abnajlae /darija-asr-corpus Darija ASR Corpus (dataset-core) Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). This repo contains four source subsets: DODa, DVoice, Wiki, and YouTube. Each subset carries its own upstream license/terms -- see below -- because they are drawn from four different original projects. Subsets Config Rows Audio bundled? Upstream license Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.audioautomatic-speech-recognition10K<n<100K0 likes539 downloads16d agoHugging Face02Williamsanderson /MedQA-Darija-MultiLingual MedQA-Darija-MultiLingual The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija. A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region. Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.audioquestion-answering100K<n<1M4 likes476 downloads5mo agoHugging Face03ohsn /darija_yt_2026 darija_yt_2026 Partition upload generated automatically. Namespace: ohsn Repo: ohsn/darija_yt_2026 Video count: 3511 Duration hours: 1565.31 This dataset contains raw audio files and per-video metadata generated from the Darija YouTube extraction pipeline. audioautomatic-speech-recognition1K<n<10K0 likes271 downloads21d agoHugging Face04adiren7 /darija_speech_to_textaudioautomatic-speech-recognition10K<n<100K13 likes180 downloads2y agoHugging Face05ayoubkirouane /darija-stt-mixThe Darija Speech To Text Dataset is a comprehensive collection designed to support speech recognition tasks for the Darija dialect, it includes audio data totaling 8.23 GB and consists of 13,178 rows of transcribed speech. This dataset covers a variety of dialects, primarily focusing on Algerian and Moroccan Darija, and also includes slang from other Arabic-speaking countries. The data has been meticulously gathered from diverse resources to ensure a rich and varied representation of spoken… See the full description on the dataset page: https://huggingface.co/datasets/ayoubkirouane/darija-stt-mix.audioautomatic-speech-recognition10K<n<100K6 likes154 downloads2y agoHugging Face0601Yassine /darija-asr-3h Moroccan Darija ASR — 3 hours YouTube Moroccan Darija, segmented and filtered, labeled with Gemini 2.5 Pro. split hours clips train 3.00 1778 validation 0.15 91 silver 0.35 184 Splits are channel-disjoint: silver channels do not appear in train. A same-size random split leaks 100% of silver channels into train. Columns id, audio (16 kHz), text (Gemini 2.5 Pro) channel (YouTube handle) duration, pesq_hyp (SQUIM, no-reference), num_speakers… See the full description on the dataset page: https://huggingface.co/datasets/01Yassine/darija-asr-3h.audioautomatic-speech-recognition1K<n<10K1 likes141 downloads24d agoHugging Face07BrunoHays /darija-speech-to-text Speech To Text Darija dataset Reupload of adiren7/darija_speech_to_text audioautomatic-speech-recognition1K<n<10K6 likes83 downloads2y agoHugging Face08abnajlae /darija-asr-benchmark-6speaker Darija ASR 6-Speaker Benchmark A fixed, paired 20-utterance benchmark read identically by 6 held-out speakers (3 female: F1, F2, F3; 3 male: M1, M2, M3 -- none present in any training corpus), used to evaluate cross-speaker generalization for a Moroccan Darija (Arabizi) Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). Consent and anonymization Written informed consent was obtained from all six speakers for the recording and… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-benchmark-6speaker.audioautomatic-speech-recognitionn<1K0 likes68 downloads16d agoHugging Face09mohamedmou /DATASET-darija Darija ASR Dataset Dataset de reconnaissance automatique de la parole en Darija marocaine. Description Langue: Darija marocaine (ary) Tache: automatic speech recognition Audio: WAV mono 16 kHz stocke en Parquet Colonnes: audio, sentence Structure Colonne Type Description audio Audio Segment audio WAV mono 16 kHz sentence string Transcription en darija License CC BY 4.0 audioautomatic-speech-recognitionn<1K1 likes46 downloads5mo agoHugging Face10anaszil /Segmented-Moroccan-Darija-Wiki-Audio-Dataset Dataset Card for Segmented Moroccan Darija Wiki Dataset Dataset Summary This dataset provides short Moroccan Darija (Moroccan Arabic) speech segments derived from the atlasia/Moroccan-Darija-Wiki-Audio-Dataset.It is a cleaned and segmented version of the parent dataset, text-cleaned with Gemini 2.5-flash and processed using a fine-tuned Whisper model for Darija (to be open-sourced soon). Each audio is split into segments of up to 30 seconds to make it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/anaszil/Segmented-Moroccan-Darija-Wiki-Audio-Dataset.audioautomatic-speech-recognition1K<n<10K1 likes39 downloads1y agoHugging Face11mohamedmou /DATASET-darija-ASR-clean Darija ASR Dataset Dataset de reconnaissance automatique de la parole en Darija marocaine. Description Langue: Darija marocaine (ary) Tache: automatic speech recognition Audio: WAV mono 16 kHz stocke en Parquet Colonnes: audio, sentence Structure Colonne Type Description audio Audio Segment audio WAV mono 16 kHz sentence string Transcription en darija License CC BY 4.0 audioautomatic-speech-recognitionn<1K0 likes30 downloads5mo agoHugging Face12BrunoHays /wikitongues-darija Wikitongues-Darija This is a small test dataset for Automatic Speech Recognition in Darija language, built from 2 captioned videos of the WikiTongues project: nawal anass Process: each webm video has been converted to monochannel 16khz wav files with ffmpeg : ffmpeg -i WIKITONGUES-_Nawal_speaking_Moroccan_Arabic.webm.1080p.vp9.webm -ar 16000 -ac 1 nawal.wav each audio has been cut in samples of less than 30 seconds audio according to the captions timestamps. The script may be… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/wikitongues-darija.audioautomatic-speech-recognitionn<1K1 likes22 downloads1y agoHugging Face13DrIAmed /darija-youtube-dataset Darija YouTube Dataset A dataset of Moroccan Arabic (Darija) speech scraped from YouTube, with transcriptions in Arabic script, Latin script (Arabizi), and English translations. Dataset Description This dataset contains sentence-level audio segments of Darija speech, paired with: Arabic transcription (modern Moroccan Arabic script) Latin transliteration (Arabizi format: 3=ع, 7=ح, 9=ق, etc.) English translation Columns Column Type Description audio… See the full description on the dataset page: https://huggingface.co/datasets/DrIAmed/darija-youtube-dataset.audioautomatic-speech-recognition1K<n<10K0 likes20 downloads7mo agoHugging Face14H20-sys /Segmented-Moroccan-Darija-Wiki-Audio-Dataset Dataset Card for Segmented Moroccan Darija Wiki Dataset Dataset Summary This dataset provides short Moroccan Darija (Moroccan Arabic) speech segments derived from the atlasia/Moroccan-Darija-Wiki-Audio-Dataset.It is a cleaned and segmented version of the parent dataset, text-cleaned with Gemini 2.5-flash and processed using a fine-tuned Whisper model for Darija (to be open-sourced soon). Each audio is split into segments of up to 30 seconds to make it suitable… See the full description on the dataset page: https://huggingface.co/datasets/H20-sys/Segmented-Moroccan-Darija-Wiki-Audio-Dataset.audioautomatic-speech-recognition1K<n<10K0 likes18 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.