CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Aregay01 /structured-wikipedia Dataset Card for Wikimedia Structured Wikipedia Quick Links Wikimedia Enterprise Structured Contents Documentation Data Dictionary Wikimedia Attribution Framework Meta-Wiki Discussion Dataset Summary Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API. This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/Aregay01/structured-wikipedia.text10M<n<100M0 likes224 downloads4mo agoHugging Face02SGzK /bridges2-aregmi1-outputtext10K<n<100K0 likes190 downloads17d agoHugging Face03Dr-AliGomaa /ar-eg-dataset ar-eg-dataset 40 h train / 10 h validation of Egyptian Arabic from a single expert speaker — Prof. Ali Gomaa, former Grand Mufti of Egypt (2003–2013), Professor of Islamic Jurisprudence at Al-Azhar University and member of Al-Azhar's Council of Senior Scholars — transcribed from his public lectures and released with his express permission. YouTube: https://www.youtube.com/@DrAliGomaa · Facebook: https://www.facebook.com/DrAliGomaa Paper: A Quran and Hadith Speech Resource and… See the full description on the dataset page: https://huggingface.co/datasets/Dr-AliGomaa/ar-eg-dataset.audioautomatic-speech-recognition1K<n<10K1 likes181 downloads1mo agoHugging Face04Aregay01 /AlphaFoldDB AlphaFoldDB Prediction Index AlphaFoldDB is an open database of predicted protein 3D structures with confidence scores, massively expanding structural coverage for known protein sequences. Splits Split Rows Parquet files train 222,017,452 12 test 24,672,064 2 total 246,689,516 14 The split is deterministic: hash(uniprot_accession) % 10 == 0 goes to test; buckets 1 through 9 go to train. Dataset Statistics Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/Aregay01/AlphaFoldDB.tabular100M<n<1B0 likes154 downloads4mo agoHugging Face05MahmoudIbrahim /ar-eg-speech-tts-multi-speakersaudio10K<n<100K0 likes133 downloads11d agoHugging Face06Aregay01 /ti-audio-feature-whisper-smalltext10K<n<100K0 likes81 downloads7mo agoHugging Face07Aregay01 /tir-fidel-whisper-small-datasettext10K<n<100K0 likes79 downloads3mo agoHugging Face08Aregay01 /tir-audio-feature-whisper-large-v3text10K<n<100K0 likes53 downloads7mo agoHugging Face09MahmoudIbrahim /tts-ar-egyption-denoisedaudio1K<n<10K1 likes39 downloads13d agoHugging Face10MohamedRashad /fleurs-ar-eg Dataset Card for FLEURS Arabic–Egyptian Edition Dataset Summary FLEURS Arabic–Egyptian Edition is an unofficial, language-specific subset and adaptation of the FLEURS dataset, focused on Arabic (Egyptian) speech data. The dataset is designed for Automatic Speech Recognition (ASR) research and evaluation and follows the original FLEURS structure while being packaged as a standalone Arabic-focused dataset. The data originates from the FLEURS (Few-shot Learning Evaluation of… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/fleurs-ar-eg.audioautomatic-speech-recognition1K<n<10K1 likes33 downloads9mo agoHugging Face11Aregay01 /ti-audio-feature-whisper-large-v3text10K<n<100K0 likes23 downloads7mo agoHugging Face12asas-ai /mlqa_parallel_ar_eg Dataset Card for "mlqa_parallel_ar_eg" More Information needed textquestion-answering1K<n<10K0 likes19 downloads2y agoHugging Face13Aregay01 /ti-audio-feature-whisper-mediumtext10K<n<100K0 likes18 downloads7mo agoHugging Face14Aregay01 /tir-audio-feature-whisper-v3text10K<n<100K0 likes11 downloads7mo agoHugging Face15Aregay01 /tir-audio-feature-whisper-mediumtext10K<n<100K0 likes7 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.