moganai/mogan-turkish-web
Mogan Turkish Web A large-scale Turkish web corpus derived from monthly Common Crawl snapshots covering the period from January 2025 to June 2026. The corpus was constructed by extracting Turkish-language content from raw Common Crawl WARC/WET dumps, followed by language filtering, PII masking, and near-duplicate removal. 📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum Dataset Summary This dataset… See the full description on the dataset page: https://huggingface.co/datasets/moganai/mogan-turkish-web.
Mogan Turkish Web
<p align="center"> <a href="https://buymeacoffee.com/moganai"><img src="https://img.shields.io/badge/BuyMea_Coffee-FFDD00?style=flat-square&logo=buymeacoffee&logoColor=black" alt="Buy Me a Coffee"/></a> </p>
A large-scale Turkish web corpus derived from monthly Common Crawl snapshots covering the period from January 2025 to June 2026. The corpus was constructed by extracting Turkish-language content from raw Common Crawl WARC/WET dumps, followed by language filtering, PII masking, and near-duplicate removal.
📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum
Dataset Summary
This dataset contains 32,462,870 documents and roughly 23 billion tokens, drawn from monthly Common Crawl acquisitions.
Data Collection and Processing
- Acquisition: Turkish-language pages were extracted from monthly Common Crawl snapshots (
commoncrawl_awssource). - Language filtering: each document carries a
lang_scorereflecting confidence in Turkish-language identification; low-confidence documents were excluded upstream. - PII masking: documents were scanned for personally identifiable information (national ID numbers, phone numbers, email addresses, IBANs, and card numbers); detected instances were masked prior to release. The
had_pii_maskedfield indicates whether any masking was applied to a given document. - Near-duplicate removal: a MinHash + LSH pipeline (126 permutations, 14 bands of 9 rows each, approximate Jaccard similarity threshold of 0.75) was applied to remove near-duplicate documents.
Data Fields
Intended Use
This dataset is intended for pretraining or continued pretraining of Turkish (or multilingual) language models. As with any web-derived corpus, it may contain noisy, low-quality, or offensive content despite the filtering steps described above; downstream users should apply additional quality or safety filtering appropriate to their use case.
License
Released under ODC-BY. Common Crawl data is subject to Common Crawl's Terms of Use.
Citation
@misc{moganturkishweb2026,
title={Mogan Turkish Web},
author={Furkan Yılmaz and Aleyna Taşdemir and Faruk Gözay and Bekir Top},
year={2026},
publisher={Hugging Face},
howpublished={\url{https://huggingface.co/datasets/moganai/mogan-turkish-web}}
}Mogan Turkish Web (Türkçe)
Ocak 2025 ile Haziran 2026 arasındaki dönemi kapsayan aylık Common Crawl anlık görüntülerinden türetilmiş büyük ölçekli bir Türkçe web korpüsü. Bu korpüs, ham Common Crawl WARC/WET dökümlerinden Türkçe içeriğin çıkarılmasının ardından dil filtreleme, kişisel veri (PII) maskeleme ve yakın-kopya (near-duplicate) temizliği adımlarından geçirilerek oluşturulmuştur.
📄 Makale: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum
Veri Kümesi Özeti
Bu veri kümesi, aylık Common Crawl toplamalarından elde edilmiş 32.462.870 belge ve kabaca 23 milyar token içermektedir.
Veri Toplama ve İşleme Süreci
- Toplama: Türkçe sayfalar, aylık Common Crawl anlık görüntülerinden (
commoncrawl_awskaynağı) çıkarılmıştır. - Dil filtreleme: her belge, Türkçe dil tespitindeki güven düzeyini yansıtan bir
lang_scoredeğeri taşır; düşük güvenilirlikli belgeler önceki aşamada elenmiştir. - PII maskeleme: belgeler kişisel veri (T.C. kimlik numarası, telefon numarası, e-posta adresi, IBAN ve kart numarası) açısından taranmış, tespit edilen bilgiler yayınlanmadan önce maskelenmiştir.
had_pii_maskedalanı, ilgili belgeye herhangi bir maskeleme uygulanıp uygulanmadığını belirtir. - Yakın-kopya temizliği: yakın-kopya belgeleri kaldırmak için bir MinHash + LSH pipeline'ı (126 permütasyon, 14 bant x 9 satır, ~0.75 Jaccard benzerlik eşiği) uygulanmıştır.
Alan Şeması
Kullanım Amacı
Bu veri kümesi, Türkçe (veya çok dilli) dil modellerinin ön-eğitimi ya da sürdürülen ön-eğitimi için tasarlanmıştır. Web kökenli her korpüste olduğu gibi, yukarıda açıklanan filtreleme adımlarına rağmen gürültülü, düşük kaliteli veya sakıncalı içerik barındırabilir; alt görevlerde çalışan kullanıcıların kendi kullanım senaryolarına uygun ek kalite veya güvenlik filtrelemesi uygulaması önerilir.
Lisans
ODC-BY lisansı ile yayınlanmıştır. Common Crawl verisi, Common Crawl'ın Kullanım Koşullarına tabidir.
Atıf
@misc{moganturkishweb2026,
title={Mogan Turkish Web},
author={Furkan Yılmaz and Aleyna Taşdemir and Faruk Gözay and Bekir Top},
year={2026},
publisher={Hugging Face},
howpublished={\url{https://huggingface.co/datasets/moganai/mogan-turkish-web}}
}<p align="center"> <b>Support MoganAI</b><br/> If our open Turkish models and datasets are useful to you, you can support our work.<br/> Açık Türkçe modellerimiz ve veri setlerimiz işinize yarıyorsa çalışmalarımıza destek olabilirsiniz. </p>
<p align="center"> <a href="https://buymeacoffee.com/moganai"><img src="https://cdn.buymeacoffee.com/buttons/v2/default-yellow.png" alt="Buy Me a Coffee" width="300"/></a> </p>
