datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-llm-dataset
Turkish Pretraining Corpus
Dataset Description
This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models.
This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/tascib/turkish-llm-dataset.turkish-raw-text-cleaned
Turkish Raw Text Cleaned
turkish-raw-text-cleaned, Türkçe dil modeli çalışmaları için hazırlanmış temizlenmiş ham metin veri kümesidir. Veri kümesi, turkish-nlp-suite çatısı altında yayımlanan Türkçe metin kaynaklarının temizlenmesi, filtrelenmesi ve model eğitimine daha uygun hale getirilmesiyle oluşturulmuştur.
Bu çalışma özellikle Türkçe LLM ön-eğitimi, continual pre-training (CPT), tokenizer analizi, embedding modeli eğitimi, alan bağımsız Türkçe metin modelleme ve veri… See the full description on the dataset page: https://huggingface.co/datasets/serda-dev/turkish-raw-text-cleaned.risale-sohbet-turkish-2fineweb-2-turkish-categorized
What is this
THis is the categorized version of the Turkish subset of the fineweb-2 dataset.
It is an ongoing effort, and the details will be added soon with the rest of the dataset.
BellaTurca
Dataset Card for BellaTurca
BellaTurca is the first large-scale Turkish corpus collection for training Turkish language models. The total size is around 245GB and 30 billion words. BellaTurca's focus is high quality, diversity as well as the size.
This collection is made up of five datasets: AkademikDerlem, OzenliDerlem, ForumSohbetleri, Temiz OSCAR and Temiz mC4. Originally there was a book corpus included, but it is excluded due to containing copyrighted material.
AkademikDerlem… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/BellaTurca.turkishfineweb2-cleaned
TurkishFineweb2-Cleaned
A Turkish web corpus derived from the Turkish (tur_Latn) subset of
FineWeb-2, augmented with an additional quality-classification layer
and a near-duplicate removal pass.
📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum
Source
FineWeb-2 is a
large-scale, multilingual web corpus built from Common Crawl. This dataset
covers the Turkish (tur_Latn) portion of FineWeb-2, spanning the… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkishfineweb2-cleaned.Turkish-Python-instruction
🚀 DİKKAT VERİ SETİ GÜNCELLENME SÜRECİNE ALINMIŞTIR LÜTFEN AÇIKLAMAYI OKUYUNUZ. Turkish Python & System Engineering Dataset (BYSISMO v2.0)
25 Kategorilik Büyük Türkçe Python & Sistem Mühendisliği Havuzu
📢 SÜRÜM & DOĞRULAMA DURUMU (VERSION ROADMAP)
v1.0 (Eski Arşiv - 289K / 8 Kategori): Yüksek kalite standartlarımız gereği yeniden yapılandırmaya alınmış ve dondurulmuştur.
v2.0 (Yeni Master Sürüm - 416K+ / 17 Kategori): Kodlar yalnızca sözdizimi… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/Turkish-Python-instruction.Turkish_corpus
Turkish Corpus 🇹🇷
Turkish Corpus is a large-scale cleaned Turkish text dataset created by collecting public Turkish corpora and extracting Turkish-language portions from multilingual datasets. The dataset is designed for Turkish Natural Language Processing research, language model pretraining, tokenizer training, embedding models, retrieval systems, and general Turkish language understanding tasks.
The main purpose of this dataset is to provide a practical, scalable, and… See the full description on the dataset page: https://huggingface.co/datasets/Ethosoft/Turkish_corpus.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/mrfg/turkish-court-decisions.turkish_embedding_model_training_datamogan-turkish-web
Mogan Turkish Web
A large-scale Turkish web corpus derived from monthly Common Crawl snapshots
covering the period from January 2025 to June 2026. The corpus was
constructed by extracting Turkish-language content from raw Common Crawl
WARC/WET dumps, followed by language filtering, PII masking, and
near-duplicate removal.
📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum
Dataset Summary
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/moganai/mogan-turkish-web.turkish-corpus-100b
Turkish Corpus 100B (TC-100B)
Dataset Summary
The Turkish Corpus 100B (TC-100B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Turkish. Comprising approximately 105 Billion tokens (measured with Qwen/Llama3 tokenizer), it represents one of the largest open resources for Turkish LLM pretraining.
The dataset is engineered for a two-stage training pipeline:
Pretrain Subset (~103B Tokens): A diverse mix of… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/turkish-corpus-100b.turkish-tts-combined-raw
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından düzenlenmiştir. Orijinal veri seti afkfatih tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: afkfatih/turkish-tts-combined-raw
🔗 Derleyen Platform: VeriPazarı
Türkçe TTS Birleşik Veri Seti (Turkish TTS Combined)
7 farklı açık kaynak Türkçe TTS (Metinden Sese) veri setinin birleşimidir.
~81.500 örnek |… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/turkish-tts-combined-raw.turkish_instructionsOzenliDerlem
Dataset Card for OzenliDerlem
OzenliDerlem (a.k.a CraftedCrawl) is a carefully assembled collection of top-notch web crawl data from handpicked websites, featuring articles, journals, and magazines. It focuses on gathering rich and detailed text content, especially longer articles. The collection covers a wide range of topics, including travel, news, culture, fairy tales and folklore, movie reviews, popular science, product and service complaints, fashion and self-care, trendy… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/OzenliDerlem.turkish-sentiment-analysis-dataset
Dataset
This dataset contains positive , negative and notr sentences from several data sources given in the references. In the most sentiment models , there are only two labels; positive and negative. However , user input can be totally notr sentence. For such cases there were no data I could find. Therefore I created this dataset with 3 class. Positive and negative sentences are listed below. Notr examples are extraced from turkish wiki dump. In addition, added some random text… See the full description on the dataset page: https://huggingface.co/datasets/winvoker/turkish-sentiment-analysis-dataset.temiz-mC4
Dataset Card for Temiz mC4
Temiz mC4 is the cleaned version of CulturaX corpus' Turkish split.
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
split
num instances
size
num of words
train
76.432.893
168GB
21.06B
Total
76.432.893
168GB
21.06B
This collection includes web text, crawled from internet for mC4 corpus. CulturaX is even refined version of mC4 with quality filtering and… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-mC4.TrGLUE
TrGLUE - The First Non-Translate Natural Language Understanding Benchmark for Turkish
Dataset Card for TrGLUE
TrGLUE is a natural language understanding benchmarking dataset including several single sentence and sentence pair classification tasks.
The inspiration is clearly the original GLUE benchmark.
Tasks
Single Sentence Tasks
TrCOLA The original Corpus of Linguistic Acceptability consists of sentences compiled from English literature textbooks.… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/TrGLUE.fineweb2-hq-turkishCosmos-Turkish-Corpus-v1.0This is the Turkish pretraining corpus of the Cosmos AI Research Group.
It contains ~15B tokens and demonstrates competitive performance across various Turkish benchmarks when used in continual pretraining setups.
Cosmos-Turkish-Corpus is collected from a wide range of Turkish websites, including forums, news sources, Wikipedia, and more.
URL-based deduplication has been applied; however, additional content-level deduplication and filtering may be required before use.
turkish-daily-dialogues-5k
Turkish Daily Dialogues 5K
Exactly 5,000 synthetic, multi-turn Turkish conversations covering ordinary daily-life situations. The corpus is designed as a small, auditable baseline for dialogue modelling, instruction-format experiments, augmentation research, and Turkish-language evaluation—not as a substitute for conversations written by real people.
Provenance in one sentence: the Turkish source scenario library was drafted with AI assistance specifically for this project… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/turkish-daily-dialogues-5k.image-captioning-turkish
Türkçe Image Captioning Veri Seti
Bu veri seti BLIP3o modelinin pretrain eğitiminde kullanılan BLIP3o-Pretrain-Long-Caption ve BLIP3o-Pretrain-Short-Caption veri setlerinin Türkçeye çevirilmiş bir alt parçasıdır. Orijinal veri setinin oluşturulması ile ilgili detaylı bilgiye BLIP-3o makalesi üzerinden ulaşabilirsiniz.
Veri seti Image-to-Text modellerinin eğitilmesinde veya ince ayar sürecinde kullanılabilir. Veri seti, orijinal veri setinin lisansı olan Apache 2.0 altında… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/image-captioning-turkish.turkish_corpusGoogle-Translated_Turkish_GPQA_Datasetcleaned_turkish_embedding_model_training_data_colabtemiz-OSCAR
Dataset Card for Temiz OSCAR
Temiz OSCAR is a corpora collection consisting of cleaned versions of original OSCAR corpora.
This collection is made up of four datasets: OSCAR-2019, OSCAR-2109, OSCAR-2201 and OSCAR-2301
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
Dataset
num instances
size
num of words
OSCAR-2019
3.671.430
7.7G
976M
OSCAR-2109
8.472.809
18G
2.22B
OSCAR-2201… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-OSCAR.Turkish_TTS_DataTurkish-SFT-Dataset-v1.0
Turkish-SFT-Dataset-v1.01
Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci
🔎 Özet
Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.Turkish-AlpacaStanford alpaca turkish: Stanford Alpaca
cleaned_turkish_embedding_model_training_data_colab
Citation
If you use this dataset in your research, please cite the following paper:
@inproceedings{baysan-gungor-2025-tr,
title = "{TR}-{MTEB}: A Comprehensive Benchmark and Embedding Model Suite for {T}urkish Sentence Representations",
author = "Baysan, Mehmet Selman and
Gungor, Tunga",
booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2025",
month = nov,
year = "2025",
address = "Suzhou, China",
publisher =… See the full description on the dataset page: https://huggingface.co/datasets/trmteb/cleaned_turkish_embedding_model_training_data_colab.
