MohamedGomaa30/EGYSpeak
EGYSpeak A curated dataset of 147,979 single-speaker Egyptian Arabic (pure dialect) audio clips with transcriptions, sourced from the fadisarwat/egyptian-arabic-lines Kaggle dataset and processed through a rigorous ASR pipeline. Quick Start 1. Download the dataset: from huggingface_hub import snapshot_download snapshot_download( repo_id="MohamedGomaa30/EGYSpeak", repo_type="dataset", local_dir="EGYSpeak", ) 2. Extract the dataset: from… See the full description on the dataset page: https://huggingface.co/datasets/MohamedGomaa30/EGYSpeak.
Update README.md
Update README.md
Upload README.md with huggingface_hub
Upload egyspeak_loader.py with huggingface_hub
Update README.md
Rename train.csv to metadata.csv
Rename train.jsonl to metadata.jsonl
Upload README.md with huggingface_hub
Delete EGYSpeak.py with huggingface_hub
Rename metadata.jsonl to train.jsonl
Rename metadata.csv to train.csv
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload EGYSpeak.py with huggingface_hub
Add files using upload-large-folder tool
initial commit
