CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01karimBD /saudi1audio100K<n<1M0 likes302 downloads5mo agoHugging Face02Mans1611 /Lahgtna-saudi Lahgtna Saudi (Mans1611/Lahgtna-saudi) Saudi-dialect subset prepared for Arabic ASR fine-tuning (from oddadmix/dialectal-arabic-lahgtna-v2, filtered to language == "sa"). Splits Split Rows train 11,030 test 581 Columns audio, text, language, duration Features {'audio': Audio(sampling_rate=16000, decode=False, num_channels=None, stream_index=None), 'text': Value('string'), 'language': Value('string'), 'duration':… See the full description on the dataset page: https://huggingface.co/datasets/Mans1611/Lahgtna-saudi.audioautomatic-speech-recognition10K<n<100K0 likes235 downloads1mo agoHugging Face03karimBD /saudi2audio100K<n<1M0 likes223 downloads5mo agoHugging Face04Rabe3 /saudi-english-code-switching-datasetaudio10K<n<100K0 likes221 downloads7mo agoHugging Face05HeshamHaroon /arabic-msa-25k-saudi-male-tashkeel Arabic MSA 25K — Saudi Male (Tashkeel) 25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories. Dataset Summary arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA) speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.tabulartext-to-speech10K<n<100K10 likes186 downloads5mo agoHugging Face06musabalosimi /saudi_dialect_asrv1.0audio1K<n<10K4 likes179 downloads1y agoHugging Face07karimBD /saudi2_conaudio10K<n<100K0 likes176 downloads5mo agoHugging Face08AhmedEladl /saudi-dialect-speech-female 🌍 Saudi Dialectal Arabic Audio Dataset This repository contains cleaned, segmented, and dual-transcribed Arabic speech data intended for speech modeling, ASR benchmarking, and Text-to-Speech (TTS) fine-tuning. 🗂️ Dataset Columns Column Description audio The audio chunk (22,050 Hz, mono WAV) duration Chunk duration in seconds base_transcription Transcript from the base Arabic ASR model dialectal_transcription Transcript from the Saudi-dialectal… See the full description on the dataset page: https://huggingface.co/datasets/AhmedEladl/saudi-dialect-speech-female.audioautomatic-speech-recognition1K<n<10K1 likes154 downloads1mo agoHugging Face09karimBD /saudi1_con_tempaudio10K<n<100K0 likes150 downloads5mo agoHugging Face10Rabe3 /saudi-tts-synthetic-200kaudio100K<n<1M0 likes149 downloads3mo agoHugging Face11shinnnsh23 /30k-SADA22_Saudiaudio10K<n<100K1 likes117 downloads2mo agoHugging Face12AhmedEladl /saudi-dialect-speech-maleaudio1K<n<10K0 likes103 downloads1y agoHugging Face13Nexdata-AI /268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Data-Sample Description Arabic(Saudi) Multi-stream Spontaneous Dialogue Smartphone speech dataset-Customer Service. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers(268 native speakers), geographicly speaking, enhancing model performance in real and complex tasks. For more details, please refer to the link: https://www.nexdata.ai/datasets/speechrecog/1627?source=Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Data-Sample.audion<1K0 likes87 downloads1mo agoHugging Face14Abdelrahman2922 /arabic-tts-saudi-multi-speaker-xtts Arabic Saudi TTS Dataset (LJSpeech Format) 🇸🇦 This dataset is designed for training Text-to-Speech (TTS) models such as XTTS_v2 using the LJSpeech format. 📌 Overview Language: Arabic (Saudi Dialect) Format: LJSpeech Use Case: TTS training (XTTS_v2, YourTTS, Tacotron, etc.) Speakers: Multi-speaker (Male & Female) Audio Format: WAV (mono recommended) Sample Rate: 22050 Hz (recommended) 📂 Structure all_data/ │ ├── wavs/ │ ├── sample_0.wav │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Abdelrahman2922/arabic-tts-saudi-multi-speaker-xtts.audio1K<n<10K2 likes68 downloads6mo agoHugging Face15Nexdata-kr /268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Dataset Description 사우디아라비아 아랍어(Arabic-Saudi) 멀티스트림 자연 대화 스마트폰 고객 서비스 음성 데이터셋입니다. 다양한 고객 서비스 상황에서 자유롭게 대화하는 방식으로 수집되었으며, 전사 텍스트와 함께 화자 ID, 성별, 연령 등의 속성 정보가 제공됩니다. 총 268명의 아랍어 원어민 화자로부터 데이터를 수집하여 다양한 화자 특성을 반영했으며, 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상을 지원합니다. 자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1627?source=hf.kr Specifications Format 16kHz, 16 bit, WAV, 모노 채널 Content category 정해진 주제 없이 자유롭게 진행된 자연 대화… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Dataset.audion<1K0 likes68 downloads13d agoHugging Face16ranaRan689 /SaudiTalk SaudiTalk: A Multi-Source Dialectal Speech Dataset from Saudi Arabia Dataset Summary SaudiTalk is a curated and human-verified Arabic speech dataset covering three major Saudi dialects: Hijazi, Ha’il, and Southern. The dataset is constructed from publicly available social media content and is designed to support research in Automatic Speech Recognition (ASR), dialect identification, and Arabic speech processing. Key Features 3 Saudi dialects:… See the full description on the dataset page: https://huggingface.co/datasets/ranaRan689/SaudiTalk.audion<1K0 likes57 downloads10d agoHugging Face17HuggingPanda /SDAIANCAI-Saudilang-Code-Switch-Corpusaudio1K<n<10K2 likes42 downloads2y agoHugging Face18HuggingPanda /Saudi_Podcasts_ASRaudio1K<n<10K6 likes39 downloads2y agoHugging Face19Tadabbur /yasser-aldosari-saudi-centeraudion<1K0 likes30 downloads2mo agoHugging Face20musabalosimi /saudi_asraudio1K<n<10K1 likes26 downloads1y agoHugging Face21Nasser2023 /saudi_arabic_accentaudion<1K0 likes25 downloads3y agoHugging Face22MahmoudIbrahim /60H-SADA22-Saudiaudio10K<n<100K0 likes25 downloads8mo agoHugging Face23AhmedAshrafMarzouk /arabic-tts-saudi-audio-datasetaudio1K<n<10K0 likes25 downloads4mo agoHugging Face24Rabe3 /saudi-tts-synthetic-100kaudio1K<n<10K0 likes20 downloads3mo agoHugging Face25DataoceanAI /Free_Dialogue_in_Saudi_Arabia_Corpus Description This dataset covers multiple scenarios such as banking, healthcare, insurance, sales, telecom, travel. The speakers are gender evenly, and each set of the audio is approximately 0.5 hour. For more details, please refer to the link: https://dataoceanai.com/datasets/asr/free-dialogue-in-saudi-arabia-corpus/ Specification ID: King-ASR-919 SIZE: 113 hours LANGUAGE: Saudi Arabia SPEAKERS: 100 AGES: 18-45 years old DEVICES: Mobile audion<1K0 likes19 downloads2y agoHugging Face26AhmedAshrafMarzouk /saudi-podcast-1hraudion<1K0 likes10 downloads4mo agoHugging Face27ismailelsayedeltanja /Saudi-data-ttsaudion<1K0 likes9 downloads2mo agoHugging Face28imaneek /saudi-cs-datasetgatedaudion<1K0 likes3 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.