datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BellaTurca
Dataset Card for BellaTurca
BellaTurca is the first large-scale Turkish corpus collection for training Turkish language models. The total size is around 245GB and 30 billion words. BellaTurca's focus is high quality, diversity as well as the size.
This collection is made up of five datasets: AkademikDerlem, OzenliDerlem, ForumSohbetleri, Temiz OSCAR and Temiz mC4. Originally there was a book corpus included, but it is excluded due to containing copyrighted material.
AkademikDerlem… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/BellaTurca.OzenliDerlem
Dataset Card for OzenliDerlem
OzenliDerlem (a.k.a CraftedCrawl) is a carefully assembled collection of top-notch web crawl data from handpicked websites, featuring articles, journals, and magazines. It focuses on gathering rich and detailed text content, especially longer articles. The collection covers a wide range of topics, including travel, news, culture, fairy tales and folklore, movie reviews, popular science, product and service complaints, fashion and self-care, trendy… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/OzenliDerlem.TrGLUE
TrGLUE - The First Non-Translate Natural Language Understanding Benchmark for Turkish
Dataset Card for TrGLUE
TrGLUE is a natural language understanding benchmarking dataset including several single sentence and sentence pair classification tasks.
The inspiration is clearly the original GLUE benchmark.
Tasks
Single Sentence Tasks
TrCOLA The original Corpus of Linguistic Acceptability consists of sentences compiled from English literature textbooks.… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/TrGLUE.temiz-OSCAR
Dataset Card for Temiz OSCAR
Temiz OSCAR is a corpora collection consisting of cleaned versions of original OSCAR corpora.
This collection is made up of four datasets: OSCAR-2019, OSCAR-2109, OSCAR-2201 and OSCAR-2301
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
Dataset
num instances
size
num of words
OSCAR-2019
3.671.430
7.7G
976M
OSCAR-2109
8.472.809
18G
2.22B
OSCAR-2201… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-OSCAR.AkademikDerlem
Dataset Card for AkademikDerlem
AkademikDerlem is a scientific text corpus for Turkish, gathered from misc academical publication websites.
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
This collection is made up of five datasets: Articles, Academic-Abstracts, Medical-Articles, Medical-Abstracts, and Bilkent-Writings. The Bilkent-Writings dataset comes from creative writings produced in the… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/AkademikDerlem.Havadis
Dataset Card for Havadis
Havadis is a high quality and large Turkish news corpus, indeed the largest Turkish news corpus ever.
This corpus is scraped from online news sebsites and includes text from popular newspapers such as
CNN Türk
Habertürk
Hürriyet
Millyet
NTV
Posta
Sabah
Star
Sözcü
Takvim
. The instances are first crawled from the corresponding websites, then went throught an extensive cleaning process. We eliminated instances that are too short, too repetetive… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Havadis.TurkishHateMap
Turkish Hate Map - A Large Scale and Diverse Hate Speech Dataset for Turkish
Dataset Summary
Turkish Hate Map (TuHaMa for short) is a big scale Turkish hate speech dataset that includes diverse target groups such as misogyny,
political animosity, animal aversion, vegan antipathy, ethnic group hostility, and more. The dataset includes a total of 52K instances with 13 target groups.
The dataset includes 4 labels, offensive, hate, neutral and civilized.
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/TurkishHateMap.InstrucTurca
InstrucTurca v1.0.0 is a diverse synthetic instruction tuning dataset crafted for instruction-tuning Turkish LLMs. The data is compiled data various English datasets and sources, such as code instructions, poems, summarized texts, medical texts, and more.
Dataset content
BI55/MedText
checkai/instruction-poems
garage-bAInd/Open-Platypus
Locutusque/ColumnedChatCombined
nampdn-ai/tiny-codes
Open-Orca/OpenOrca
pubmed_qa
TIGER-Lab/MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/InstrucTurca.ForumSohbetleri
Dataset Card for ForumSohbetleri
ForumSohbetleri a web forum tetx corpus for Turkish, indeed first large-scale Turkish forum text corpus.
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
This collection is made up of several subsets, each subset is gathered from the corresponding forum website. Forum websites contains diverse topics, ladies only, tech, economics, life, relations and much more...… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/ForumSohbetleri.turkish-wikiNER
Dataset Card for "turkish-nlp-suite/turkish-wikiNER"
Dataset Summary
Turkish NER dataset from Wikipedia sentences. 20.000 sentences are sampled and re-annotated from Kuzgunlar NER dataset.
Annotations are done by Co-one. Many thanks to them for their contributions. This dataset is also used in our brand new spaCy Turkish packages.
Dataset Instances
An instance of this dataset looks as follows:
{
"tokens": ["Çekimler", "5", "Temmuz", "2005", "tarihinde"… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/turkish-wikiNER.eksisozluk-ekonomi-ve-finans-tr
Ekşi Sözlük Türkçe Teknoloji Dataset
Ekşi Sözlük'ten derlenen, teknoloji kategorisine ait Türkçe kullanıcı entry'lerinden oluşan bir veri setidir. Türkçe NLP araştırmaları ve LLM eğitimi için hazırlanmıştır.
İçerik
Makroekonomik kavramlar, yatırım araçları, kripto para ve güncel ekonomik gelişmeleri kapsayan 39 farklı başlık altında toplanmış entry'lerden oluşmaktadır.
Alan
Kapsanan Başlıklar
Ekonomik Kriz & Genel Durum
2025-2026 ekonomik krizi, devalüasyon… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-datasets/eksisozluk-ekonomi-ve-finans-tr.SentiTurca
SentiTurca - A Sentiment Analysis Benchmark for Turkish
Dataset Card for SentiTurca
SentiTurca is a sentiment analysis benchmarking dataset including movie reviews, hate speech and e-commerce reviews classification.
Datasets
e-commerce: The e-commerce reviews are scraped from e-commerce websites Trendyol.com and Hepsiburada.com, including review for many product types such as cloths, toys, books, electronics and more.E-commerce reviews has their stand… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/SentiTurca.turkish-morph-analysis
Dataset Card for TrMorphTester
This dataset is a testing dataset for Turkish morphology, aiming to calculate how other subword strategies aligns with morphological segmentation of Turkish.
The data is automatically generated from Turkish morpoholigical lexicon. For each row, we offer a surface form, then lemma and all suffixes, separated by a + character.
The dataset has several splits for several purposes:
lemma: Surface form is same with lemma, no suffixes at all. Nouns, verbs… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/turkish-morph-analysis.vitamins-supplements-reviews
Dataset Card for turkish-nlp-suite/vitamins-supplements-reviews
Dataset Summary
Turkish sentiment analysis dataset from customer reviews about supplement and vitamin products. The dataset is scraped from Vitaminler.com and contains
customer reviews and star rating about vitamin and supplement products.
Each customer review in the Vitamins and Supplements Reviews Dataset describes a customer’s experience with a supplement product in terms of the product’s effectiveness… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/vitamins-supplements-reviews.temiz-WikiA cleaned version of Turkish Wikipedia dataset. The soource is wikimedia/wikipedia repo.
The text is cleaned throughoutly, first of all we eliminated text that are shorter than a predetermined threshold of words and characters. Then we normalized with NFKC, cleaned some non-ASCII chars and normalized whitespaces.
beyazperde-top-300-movie-reviews
Dataset Card for turkish-nlp-suite/beyazperde-top-300-movie-reviews
Dataset Summary
Beyazperde Movie Reviews offers Turkish sentiment analysis datasets that is scraped from popular movie reviews website Beyazperde.com. Top 300 Movies include audience reviews about best 300 movies of all the time. Here's the star rating distribution:
star rating
count
0.5
101
1.0
39
1.5
19
2.0
44
2.5
210
3.0
196
3.5
490
4.0
1212
4.5
818
5.0
1251
total
4380… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/beyazperde-top-300-movie-reviews.beyazperde-all-movie-reviews
Dataset Card for turkish-nlp-suite/beyazperde-all-movie-reviews
Dataset Summary
Beyazperde Movie Reviews offers Turkish sentiment analysis datasets that is scraped from popular movie reviews website Beyazperde.com. All Movie Reviews include audience reviews about movies of all the time. Here's the star rating distribution:
star rating
count
0.5
3.635
1.0
2.325
1.5
1.077
2.0
1.902
2.5
4.767
3.0
4.347
3.5
6.495
4.0
9.486
4.5
3.652
5.0
7.594… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/beyazperde-all-movie-reviews.MusteriYorumlari
MüşteriYorumlari - A Large Scale Customer Sentiment Analysis Dataset for Turkish
Dataset Summary
MüşteriYorumları is a Turkish e-commerce customer reviews dataset of size 103K, scraped from Hepsiburada.com and Trendyol.com. These reviews encompass a wide
array of product categories, including apparel, food items, baby products, and books. Review stars are in range of 1-5 stars.
The star distribution is as follows:
star rating
count
1
12,873
2
11,472
3
18… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/MusteriYorumlari.economy-and-finance
Ekşi Sözlük Türkçe Teknoloji Dataset
Ekşi Sözlük'ten derlenen, teknoloji kategorisine ait Türkçe kullanıcı entry'lerinden oluşan bir veri setidir. Türkçe NLP araştırmaları ve LLM eğitimi için hazırlanmıştır.
İçerik
Yapay zeka, ChatGPT, sosyal medya algoritmaları, kripto para ve diğer teknoloji konularını kapsayan 24 farklı başlık altında toplanmış entry'lerden oluşmaktadır.
Alan
Kapsanan Başlıklar
Yapay Zeka
yapay zeka, chatgpt, claude ai, meta ai, apple… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-datasets/economy-and-finance.eksisozluk_teknoloji_tr
Ekşi Sözlük Türkçe Teknoloji Dataset
Genel Bilgi
Dil: Türkçe
Kaynak: Ekşi Sözlük
Konu: Teknoloji (Yapay Zeka, ChatGPT, Sosyal Medya vb.)
Format: JSONL
Toplam Entry: 7521
Alan Açıklamaları
text: Entry içeriği
author: Yazar nickname
date: Giriş tarihi
entry_no: Tekil entry ID
topic: Başlık kategorisi
source: eksisozluk
language: tr
Kullanım Amacı
Türkçe LLM eğitimi ve fine-tuning için uygundur.
Lisans
CC BY-NC 4.0
Corona-mini
Dataset Card for turkish-nlp-suite/Corona-mini
Dataset Summary
This is a tiny Turkish corpus consisting of comments about Corona symptoms. The corpus is compiled from two Ekşisözlük headlines "covid-19 belirtileri" and "gün gün koronavirüs belirtileri":
https://eksisozluk.com/covid-19-belirtileri--6416646
https://eksisozluk.com/gun-gun-koronavirus-belirtileri--6757665
This corpus
contains 178 raw, 175 processed comments
all comments are in Turkish
comes in 2… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Corona-mini.sinefil-movie-reviews
Sinefil Movie Reviews
A movie reviews sentiment analysis dataset for Turkish.
Dataset Summary
Sinefil Movie Reviews offers Turkish sentiment analysis datasets that is scraped from popular movie reviews website Sinefil.com. The reviews include audience reviews about movies of all the time.
The score field takes values between 1 and 9.9. Values are like 8, 8.1, 8.2 .. 8.9. Here's the distribution divided into integer bins:
star rating
count
1-2
2323
2-3
874… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/sinefil-movie-reviews.TrCOLA
TrCOLA - Corpus of Linguistic Acceptability for Turkish Language
Dataset Card for TrCOLA
TrCOLA is the Turkish version of CoLA dataset, The Corpus of Linguistic Acceptability.
This dataset introduces linguistic acceptability task for Turkish. The total dataset size is 9.9K instances.
Each instance of the dataset is an original and correct sentence, variation of sentence that is produced in a specific way, the variation type and a binary label stating the sentence is a… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/TrCOLA.BuyukSinema
BüyükSinema - A Large Scale Turkish Movie Reviews Sentiment Dataset
Dataset Summary
BüyükSinema is a Turkish movie reviews dataset of size 87K, scraped from Sinefil.com and Beyazperde.com. Hence this dataset is a superset of
BeyazPerde All Movie Reviews,
BeyazPerde Top 300 Movie Reviews and
Sinefil Movie Reviews datasets.
This is a merge of the three different datasets from two resources, hence we scaled the output stars into the range of 1-10 accordingly.
The star… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/BuyukSinema.vitamins-supplements-NER
Dataset Card for turkish-nlp-suite/vitamins-supplements-NER
Dataset Summary
The Vitamins and Supplements NER Dataset is a NER dataset containing customer reviews with entity and span annotations. User reviews were collected from a popular supplement products e-
commerce website Vitaminler.com.
Each customer review in the Vitamins and Supplements NER Dataset describes a customer’s experience with a supplement product in terms of that product’s effectiveness, side… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/vitamins-supplements-NER.Treebank-Benchmarking
Turkish Treebank Benchmarking
This is the repo for Turkish treebank benchmarking, namely evaluating Tranformer models on POS-Dep-Morph task.
For the data, we used two treebank, IMST and BOUN. We converted conllu format to json lines for being compatible to HF dataset formats.
Here are treebank sizes at a glance:
Dataset
train lines
dev lines
test lines
BOUN
7803
979
979
IMST
3435
1100
1100
A typical instance from the dataset looks like:
{
"id": "ins_1267"… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Treebank-Benchmarking.turkish-org-ner
Turkish Organization Named Entity Recognition (NER) Dataset
This dataset is designed for Named Entity Recognition (NER) tasks in the Turkish language, specifically focusing on organization entities. It contains sentences annotated with the ORGANIZATION label, making it a valuable resource for training and evaluating NER models.
Usage
To use this dataset, you can load it directly from Huggingface's datasets library:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/STNM-NLPhoenix/turkish-org-ner.turkish-absa
Turkish ABSA Dataset
This repository contains a dataset created for training and evaluating an Aspect-Based Sentiment Analysis (ABSA) model in the Turkish language. The dataset includes sentences with various entities and sentiments.
About the Dataset
The dataset consists of sentences taken from customer feedback and includes tags related to the entities and sentiments mentioned in these sentences. Each review indicates which entity it is related to and the sentiment… See the full description on the dataset page: https://huggingface.co/datasets/STNM-NLPhoenix/turkish-absa.
