CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AhmedEladl /saudi-dialect-speech-female 🌍 Saudi Dialectal Arabic Audio Dataset This repository contains cleaned, segmented, and dual-transcribed Arabic speech data intended for speech modeling, ASR benchmarking, and Text-to-Speech (TTS) fine-tuning. 🗂️ Dataset Columns Column Description audio The audio chunk (22,050 Hz, mono WAV) duration Chunk duration in seconds base_transcription Transcript from the base Arabic ASR model dialectal_transcription Transcript from the Saudi-dialectal… See the full description on the dataset page: https://huggingface.co/datasets/AhmedEladl/saudi-dialect-speech-female.audioautomatic-speech-recognition1K<n<10K1 likes148 downloads1mo agoHugging Face02musabalosimi /saudi_dialect_asrv1.0audio1K<n<10K4 likes125 downloads1y agoHugging Face03AhmedEladl /saudi-dialect-speech-maleaudio1K<n<10K0 likes109 downloads1y agoHugging Face04Omartificial-Intelligence-Space /saudi-dialect-test-samples Saudi Dialect Test Samples Dataset Description This dataset contains 1280 Saudi dialect utterances across 44 categories, used for testing and evaluating the Omartificial-Intelligence-Space/SA-BERT-V1 model. The sentences represent a wide range of topics, from daily conversations to specialized domains. Dataset Structure Data Fields category: The topic category of the utterance (one of 44 categories) text: The Saudi dialect text… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/saudi-dialect-test-samples.texttext-classification1K<n<10K5 likes41 downloads1y agoHugging Face05Omartificial-Intelligence-Space /SaudiDialect-Triplet-21gated 📂 SaudiDialect-Triplet-21 : Saudi Triplet Dataset (SABER Training Data) 🧩 Dataset Summary The Saudi Triplet Dataset is a high-quality corpus of 2,964 sentence triplets (Anchor, Positive, Negative) specifically curated to capture the nuances of Saudi Arabic dialects (Najdi, Hijazi, Gulf, etc.). This dataset was created to fine-tune semantic embedding models such as SABER for tasks like Semantic Search, Retrieval-Augmented Generation (RAG), and Clustering. It covers 21… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/SaudiDialect-Triplet-21.textfeature-extraction1K<n<10K4 likes35 downloads10mo agoHugging Face06mahmoudsaalama /arabic-eou-saudi-dialect Arabic End-of-Utterance Detection Dataset (Saudi Dialect) Dataset Description This dataset is designed for training and evaluating End-of-Utterance (EOU) detection models for Arabic conversations, with emphasis on Saudi dialect patterns. Dataset Summary The dataset contains Arabic conversational samples labeled for binary classification: Positive (1): End of utterance - speaker has finished their turn Negative (0): Not end of utterance - speaker will continue… See the full description on the dataset page: https://huggingface.co/datasets/mahmoudsaalama/arabic-eou-saudi-dialect.texttext-classification1K<n<10K0 likes31 downloads10mo agoHugging Face07mahmoudsaalama /sada-eou-saudi-dialecttabular1K<n<10K0 likes13 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.