CoolFace
Datasetpublic

Prickly-Labs/1.9M-Egyptian-Corpus

1.92M Egyptian Arabic Corpus 🇪🇬 — Prickly Labs A corpus of 1.92 million Egyptian Arabic samples combining synthetic generation and real web data, built for continued pretraining and dialect adaptation of Arabic language models. Designed to reflect how Egyptians actually speak — not textbook MSA. ⚠️ This dataset contains uncensored, informal Arabic including sarcasm, humor, slang, and profanity. Use with care for public-facing applications. 📌 Overview… See the full description on the dataset page: https://huggingface.co/datasets/Prickly-Labs/1.9M-Egyptian-Corpus.

sourceHugging Facecc-by-sa-4.0updated 5mo agoView on Hugging Face
2likes131downloads

Prickly-Labs/1.9M-Egyptian-Corpus · main · files are served by the source, never re-hosted here