CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MohamedRashad /SADA22 Dataset Card for SADA (Saudi Audio Dataset for Arabic) Dataset Summary The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over 667 hours of transcribed Arabic audio recordings, primarily featuring various Saudi dialects, and was curated in a collaboration between the National Center for Artificial… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/SADA22.audioautomatic-speech-recognition100K<n<1M30 likes2.2k downloads1y agoHugging Face02mosama /sada-validation-preprocessed Details This is the SADA 2022 dataset with the input_features whish are log mels and the cleaned_labels which is the tokenized version of the cleaned_text. You can directly use this as the validation dataset when training Whisper Tiny, Small, Base & Medium models, as they all use the same tokenizer. Please double check this as well from the original model repo. In addtition, the following filters were applied to this data: All audios are less than 30 seconds and greater than 0… See the full description on the dataset page: https://huggingface.co/datasets/mosama/sada-validation-preprocessed.audioautomatic-speech-recognition1K<n<10K0 likes135 downloads1y agoHugging Face03badrex /arabic-speech-SADA22-MSA Dataset Card for SADA (Saudi Audio Dataset for Arabic) ⚠️ Caution This is only the portion of the SADA dataset where the speaker dialect is Modern Standard Arabic (MSA). To access full dataset, you should check this link. Dataset Summary The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over… See the full description on the dataset page: https://huggingface.co/datasets/badrex/arabic-speech-SADA22-MSA.audioautomatic-speech-recognition1K<n<10K2 likes88 downloads1y agoHugging Face04Ahmed007 /SADA-Najdi-from-kaggle SADA 2022 — Najdi Dialect (Preprocessed for TTS) Najdi dialect subset of SADA 2022, preprocessed and ready for XTTS-v2 fine-tuning. Preprocessing Pipeline SADA full episodes → Filter Najdi → Slice by SegmentStart/End → Resample 22050Hz → Trim silence → Peak normalize → Duration filter (2.0–11.0s) → SNR filter (≥12.0dB) → Environment filter (Clean only) → Arabic text normalization Step Details Dialect SpeakerDialect == "Najdi" Environment… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed007/SADA-Najdi-from-kaggle.audiotext-to-speech10K<n<100K0 likes60 downloads5mo agoHugging Face05badrex /arabic-speech-SADA22-Khaliji Dataset Card for SADA (Saudi Audio Dataset for Arabic) ⚠️ Caution This is only the portion of the SADA dataset where the speaker dialect is Khaliji. To access full dataset, you should check this link. Dataset Summary The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over 667 hours of… See the full description on the dataset page: https://huggingface.co/datasets/badrex/arabic-speech-SADA22-Khaliji.audioautomatic-speech-recognition10K<n<100K5 likes49 downloads1y agoHugging Face06Mohaddz /sada-diarization-previewgated SADA 2022 Arabic Diarization Training-ready speaker-attributed ASR windows derived from SADA 2022. The source recordings are mirrored at khaledalganem/sada2022. Splits train: 36,004 windows, 202.064 hours, 4,062 recordings validation: 853 windows, 4.774 hours, 88 recordings test: 901 windows, 5.006 hours, 111 recordings Total: 37,758 windows and 211.844 hours. The official SADA train, validation, and test partitions are preserved. Windows are 8–28 seconds… See the full description on the dataset page: https://huggingface.co/datasets/Mohaddz/sada-diarization-preview.audioautomatic-speech-recognition10K<n<100K0 likes14 downloads2mo agoHugging Face07Sebssihakim /Clean_One_Speaker_SADAgated Clean_One_Speaker_SADA A cleaned, single-speaker subset of the SADA (Saudi Audio Dataset for Arabic) corpus, derived from MahmoudIbrahim/100Hours-SADA22. Processing Follows the cleaning procedure from the Kaggle notebook Segmented Audio Data for Arabic Dialects (SADA): Removed rows with Unknown speaker age or gender. Removed rows whose dialect is More than 1 speaker, Unknown, or Notapplicable (every remaining segment has exactly one identified speaker).… See the full description on the dataset page: https://huggingface.co/datasets/Sebssihakim/Clean_One_Speaker_SADA.audioautomatic-speech-recognition10K<n<100K0 likes11 downloads3mo agoHugging Face08Sebssihakim /Clean_One_Speaker_SADA_Splitgated Clean_One_Speaker_SADA_Split Train/validation/test version of Sebssihakim/Clean_One_Speaker_SADA, a cleaned single-speaker subset of the SADA corpus (SDAIA / Saudi Broadcasting Authority). Splits 80/10/10, stratified by speaker_dialect, seed 42: train ~30.8k, validation ~3.9k, test ~3.9k rows. The Maghrebi dialect (7 rows) was removed — too few samples to stratify. Splits are segment-level: the same source show/speaker may appear in more than one split. Rare… See the full description on the dataset page: https://huggingface.co/datasets/Sebssihakim/Clean_One_Speaker_SADA_Split.audioautomatic-speech-recognition10K<n<100K0 likes5 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.