CoolFace
Datasetpublic

moganai/mogan-turkish-web

Mogan Turkish Web A large-scale Turkish web corpus derived from monthly Common Crawl snapshots covering the period from January 2025 to June 2026. The corpus was constructed by extracting Turkish-language content from raw Common Crawl WARC/WET dumps, followed by language filtering, PII masking, and near-duplicate removal. 📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum Dataset Summary This dataset… See the full description on the dataset page: https://huggingface.co/datasets/moganai/mogan-turkish-web.

sourceHugging Faceodc-byupdated 2d agoView on Hugging Face
6likes1.2kdownloads

moganai/mogan-turkish-web · main · files are served by the source, never re-hosted here