datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
structured-wikipedia
Dataset Card for Wikimedia Structured Wikipedia
Quick Links
Wikimedia Enterprise
Structured Contents Documentation
Data Dictionary
Wikimedia Attribution Framework
Meta-Wiki Discussion
Dataset Summary
Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API.
This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a… See the full description on the dataset page: https://huggingface.co/datasets/Aregay01/structured-wikipedia.bridges2-aregmi1-outputar-eg-dataset
ar-eg-dataset
40 h train / 10 h validation of Egyptian Arabic from a single expert speaker —
Prof. Ali Gomaa, former Grand Mufti of Egypt (2003–2013), Professor of Islamic
Jurisprudence at Al-Azhar University and member of Al-Azhar's Council of Senior Scholars —
transcribed from his public lectures and released with his express permission.
YouTube: https://www.youtube.com/@DrAliGomaa · Facebook: https://www.facebook.com/DrAliGomaa
Paper: A Quran and Hadith Speech Resource and… See the full description on the dataset page: https://huggingface.co/datasets/Dr-AliGomaa/ar-eg-dataset.AlphaFoldDB
AlphaFoldDB Prediction Index
AlphaFoldDB is an open database of predicted protein 3D structures with confidence scores, massively expanding structural coverage for known protein sequences.
Splits
Split
Rows
Parquet files
train
222,017,452
12
test
24,672,064
2
total
246,689,516
14
The split is deterministic: hash(uniprot_accession) % 10 == 0 goes to test; buckets 1 through 9 go to train.
Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/Aregay01/AlphaFoldDB.ar-eg-speech-tts-multi-speakersti-audio-feature-whisper-smalltir-fidel-whisper-small-datasettir-audio-feature-whisper-large-v3tts-ar-egyption-denoisedfleurs-ar-eg
Dataset Card for FLEURS Arabic–Egyptian Edition
Dataset Summary
FLEURS Arabic–Egyptian Edition is an unofficial, language-specific subset and adaptation of the FLEURS dataset, focused on Arabic (Egyptian) speech data.
The dataset is designed for Automatic Speech Recognition (ASR) research and evaluation and follows the original FLEURS structure while being packaged as a standalone Arabic-focused dataset.
The data originates from the FLEURS (Few-shot Learning Evaluation of… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/fleurs-ar-eg.ti-audio-feature-whisper-large-v3mlqa_parallel_ar_eg
Dataset Card for "mlqa_parallel_ar_eg"
More Information needed
ti-audio-feature-whisper-mediumtir-audio-feature-whisper-v3tir-audio-feature-whisper-medium
