Prickly-Labs/1.9M-Egyptian-Corpus
1.92M Egyptian Arabic Corpus 🇪🇬 — Prickly Labs A corpus of 1.92 million Egyptian Arabic samples combining synthetic generation and real web data, built for continued pretraining and dialect adaptation of Arabic language models. Designed to reflect how Egyptians actually speak — not textbook MSA. ⚠️ This dataset contains uncensored, informal Arabic including sarcasm, humor, slang, and profanity. Use with care for public-facing applications. 📌 Overview… See the full description on the dataset page: https://huggingface.co/datasets/Prickly-Labs/1.9M-Egyptian-Corpus.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face