CoolFace
Datasetpublic

turkish-nlp-suite/OzenliDerlem

Dataset Card for OzenliDerlem OzenliDerlem (a.k.a CraftedCrawl) is a carefully assembled collection of top-notch web crawl data from handpicked websites, featuring articles, journals, and magazines. It focuses on gathering rich and detailed text content, especially longer articles. The collection covers a wide range of topics, including travel, news, culture, fairy tales and folklore, movie reviews, popular science, product and service complaints, fashion and self-care, trendy… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/OzenliDerlem.

sourceHugging Facecc-by-sa-4.0updated 7mo agoView on Hugging Face
11likes712downloads
Dataset Card

<img src="https://raw.githubusercontent.com/turkish-nlp-suite/.github/main/profile/ozenliderlemlogo.png" width="30%" height="30%">

Dataset Card for OzenliDerlem

OzenliDerlem (a.k.a CraftedCrawl) is a carefully assembled collection of top-notch web crawl data from handpicked websites, featuring articles, journals, and magazines. It focuses on gathering rich and detailed text content, especially longer articles. The collection covers a wide range of topics, including travel, news, culture, fairy tales and folklore, movie reviews, popular science, product and service complaints, fashion and self-care, trendy tech, pop culture, and literature.

Each sub-corpus in OzenliDerlem, except for the News sub-corpus, is gathered from at least 50 high-quality websites. This variety of content helps the models better understand the finer details of modern Turkish culture, making these sub-corpora incredibly valuable. The breakdown of topics in the CraftedCrawl collection is shown below.

The News sub-corpus, originally called "Havadis," stands out as the first large-scale Turkish news crawl ever conducted. It includes data collected from 11 major newspaper websites, such as CNN Türk, Habertürk, Hürriyet, Milliyet, NTV Haber, Posta, Sabah, Sözcü, Star, T24, and Takvim.

This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.

Datasetnum instancessizenum of words (millions)
GeziNotlari33.174173M21M
Havadis744.8682.7GB322M
KulturHaritasi133.118485M59M
MasalMasal26218.7M1M
Perdearkasi-Yorumlar36.888129M16M
PopularBilim23.42497M11M
Serzenisler23.92323M2M
SusluTrendler8.99332M3M
TeknoYazilar160.680373M46M
ViralMedya146.315510M63M
YazarinKaleminden74.52970K8M
Total1.391.2394.6GB557M

Instances

A typical instance from the dataset looks like:

{
"url": "https://www.ayagimintozuyla.net/balide-edinebileceginiz-15-essiz-deneyim.html",
"text": "Önceki Bali yazılarımda da anlatığım gibi Bali her türlü gezgine kucak açan bir destinasyon. Keyif düşkünlerine, maceracı gezginlere, kendini arayanlara, veya sadece farklı bir şeyler görmek isteyenlere kadar herkes kendine ait bir şeyler bulacak Bali'de. Bu yazımda Bali'de edinebileceğiniz tecrübeleri listeleyeceğim. Kim bilir, belki yazının sonunda Bali için bilet bakmaya başlarsınız, belli mi olur 🙂"
}

Citation

@InProceedings{10.1007/978-3-031-70563-2_16,
author="Altinok, Duygu",
editor="N{\"o}th, Elmar
and Hor{\'a}k, Ale{\v{s}}
and Sojka, Petr",
title="Bella Turca: A Large-Scale Dataset of Diverse Text Sources for Turkish Language Modeling",
booktitle="Text, Speech, and Dialogue",
year="2024",
publisher="Springer Nature Switzerland",
address="Cham",
pages="196--213",
abstract="In recent studies, it has been demonstrated that incorporating diverse training datasets enhances the overall knowledge and generalization capabilities of large-scale language models, especially in cross-domain scenarios. In line with this, we introduce Bella Turca: a comprehensive Turkish text corpus, totaling 265GB, specifically curated for training language models. Bella Turca encompasses 25 distinct subsets of 4 genre, carefully chosen to ensure diversity and high quality. While Turkish is spoken widely across three continents, it suffers from a dearth of robust data resources for language modelling. Existing transformers and language models have primarily relied on repetitive corpora such as OSCAR and/or Wiki, which lack the desired diversity. Our work aims to break free from this monotony by introducing a fresh perspective to Turkish corpora resources. To the best of our knowledge, this release marks the first instance of such a vast and diverse dataset tailored for the Turkish language. Additionally, we contribute to the community by providing the code used in the dataset's construction and cleaning, fostering collaboration and knowledge sharing.",
isbn="978-3-031-70563-2"
}

Acknowledgments

This research was supported with Cloud TPUs from Google's TPU Research Cloud (TRC).