CoolFace
Datasetpublic

nassimjp/zamai-pashto-clean-cpt

ZamAI Pashto Clean CPT This dataset is a hyper-cleaned, optimized, and fully deduplicated version of tasal9/ZamAI-Pashto-Mega-Dataset intended for Causal Language Modeling (CLM), pre-training, or fine-tuning text models in the Pashto language. 🛠️ Pipeline & Filtering Details Before processing the ~1.5 GB stream, an Internal Built-In Self-Test (BIST) was executed to verify environment I/O permissions, validate character ratio logic, and check connection… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/zamai-pashto-clean-cpt.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes29downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
nassimjp/zamai-pashto-clean-cpt · CoolFace