turkish-nlp-suite/temiz-OSCAR
Dataset Card for Temiz OSCAR Temiz OSCAR is a corpora collection consisting of cleaned versions of original OSCAR corpora. This collection is made up of four datasets: OSCAR-2019, OSCAR-2109, OSCAR-2201 and OSCAR-2301 This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. Dataset num instances size num of words OSCAR-2019 3.671.430 7.7G 976M OSCAR-2109 8.472.809 18G 2.22B OSCAR-2201… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-OSCAR.
<img src="https://raw.githubusercontent.com/turkish-nlp-suite/.github/main/profile/temiz-oscar.png" width="30%" height="30%">
Dataset Card for Temiz OSCAR
Temiz OSCAR is a corpora collection consisting of cleaned versions of original OSCAR corpora. This collection is made up of four datasets: OSCAR-2019, OSCAR-2109, OSCAR-2201 and OSCAR-2301
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
This collection includes web text, crawled from internet by OSCAR project. We cleaned the corpus by quite many criterion such as text length, character to digit ratio; as well as filtered for too much profanity, adult content and more. For all the details please visit the publication.
Instances
A typical instance from the dataset looks like:
{
"text": Türkiyenin önde gelen ilaç şirketlerinden Nobel İlaç, enfeksiyonlardan korunma yolları içerikli eğitim programı ile, \
Junior Chamber International (JCI) tarafından 2016 Uluslararası Kurumsal Sosyal Sorumluluk ödülüne layık görüldü. Yarım asrı
aşkın süredir insan sağlığının korunması ve iyileştirilmesi alanında çalışan Türkiyenin önde gelen ilaç firmalarından, %100 Türk \
sermayeli Nobel İlaç, kurumsal sosyal sorumluluk alanındaki kararlılığını bu proje ile bir kez daha gösterdi. Sağlıklı yaşam \
bilincinin erken yaşlarda eğitim yoluyla geliştirilmesi gerektiğine inanan Nobel İlaç ve gönüllü çalışanları, 7 farklı oturumda \
ilkokul çağındaki 120 çocuğa ulaşarak enfeksiyonlardan korunma yolları içerikli eğitimler verdiler. 114 ülkede 169.000 üyesi \
bulunan, toplumlarda pozitif değişime ve gelişime katkıda bulunmak için gençlerin liderlik, girişimcilik becerilerini ve sosyal \
sorumluluk bilincini geliştirme misyonunu üstlenen Junior Chamber International (JCI), Nobel İlaçın bu eğitim programını 2016 Uluslararası Kurumsal Sosyal Sorumluluk Ödülüne layık gördü. Nobel İlaç bu proje ile ülkemizin geleceğini şekillendirecek çocuklarımıza, nitelikli ve kaliteli eğitim verilmesine destek olarak, sağlıklı ve bilinçli bireyler yetişmesine katkı sağlamayı hedeflemiştir. Aynı zamanda ülkemiz çocuklarında farkındalık uyandırarak, gelecekte yapılacak benzer toplumsal projelerde aktif görev almaları için onlara rol model olmayı amaçlamıştır."Citation
@InProceedings{10.1007/978-3-031-70563-2_16,
author="Altinok, Duygu",
editor="N{\"o}th, Elmar
and Hor{\'a}k, Ale{\v{s}}
and Sojka, Petr",
title="Bella Turca: A Large-Scale Dataset of Diverse Text Sources for Turkish Language Modeling",
booktitle="Text, Speech, and Dialogue",
year="2024",
publisher="Springer Nature Switzerland",
address="Cham",
pages="196--213",
abstract="In recent studies, it has been demonstrated that incorporating diverse training datasets enhances the overall knowledge and generalization capabilities of large-scale language models, especially in cross-domain scenarios. In line with this, we introduce Bella Turca: a comprehensive Turkish text corpus, totaling 265GB, specifically curated for training language models. Bella Turca encompasses 25 distinct subsets of 4 genre, carefully chosen to ensure diversity and high quality. While Turkish is spoken widely across three continents, it suffers from a dearth of robust data resources for language modelling. Existing transformers and language models have primarily relied on repetitive corpora such as OSCAR and/or Wiki, which lack the desired diversity. Our work aims to break free from this monotony by introducing a fresh perspective to Turkish corpora resources. To the best of our knowledge, this release marks the first instance of such a vast and diverse dataset tailored for the Turkish language. Additionally, we contribute to the community by providing the code used in the dataset's construction and cleaning, fostering collaboration and knowledge sharing.",
isbn="978-3-031-70563-2"
}Acknowledgments
This research was supported with Cloud TPUs from Google's TPU Research Cloud (TRC).
