datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-raw-text-cleaned
Turkish Raw Text Cleaned
turkish-raw-text-cleaned, Türkçe dil modeli çalışmaları için hazırlanmış temizlenmiş ham metin veri kümesidir. Veri kümesi, turkish-nlp-suite çatısı altında yayımlanan Türkçe metin kaynaklarının temizlenmesi, filtrelenmesi ve model eğitimine daha uygun hale getirilmesiyle oluşturulmuştur.
Bu çalışma özellikle Türkçe LLM ön-eğitimi, continual pre-training (CPT), tokenizer analizi, embedding modeli eğitimi, alan bağımsız Türkçe metin modelleme ve veri… See the full description on the dataset page: https://huggingface.co/datasets/serda-dev/turkish-raw-text-cleaned.WangchanThaiInstruct_Multi-turn_Conversation_Dataset
WangchanThaiInstruct Multi-turn Conversation Dataset
We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language.
Citation
Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633
or BibTeX
@dataset{thammaleelakul_2024_13132633,
author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.turkishfineweb2-cleaned
TurkishFineweb2-Cleaned
A Turkish web corpus derived from the Turkish (tur_Latn) subset of
FineWeb-2, augmented with an additional quality-classification layer
and a near-duplicate removal pass.
📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum
Source
FineWeb-2 is a
large-scale, multilingual web corpus built from Common Crawl. This dataset
covers the Turkish (tur_Latn) portion of FineWeb-2, spanning the… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkishfineweb2-cleaned.TurMix
TurMix (https://arxiv.org/abs/2512.18834) is a Turkish pretraining corpus containing 168 billion tokens across 219 million documents (in the minhash subset). Rather than scraping the web again, TurMix combines five publicly available Turkish datasets, applies Turkish-specific quality filtering, and performs cross-dataset deduplication.
We train a 1.4B parameter language model through nanotron on 30 billion tokens to show that the matched subset of TurMix outperforms the… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/TurMix.Turkish-Python-instruction
🚀 DİKKAT VERİ SETİ GÜNCELLENME SÜRECİNE ALINMIŞTIR LÜTFEN AÇIKLAMAYI OKUYUNUZ. Turkish Python & System Engineering Dataset (BYSISMO v2.0)
25 Kategorilik Büyük Türkçe Python & Sistem Mühendisliği Havuzu
📢 SÜRÜM & DOĞRULAMA DURUMU (VERSION ROADMAP)
v1.0 (Eski Arşiv - 289K / 8 Kategori): Yüksek kalite standartlarımız gereği yeniden yapılandırmaya alınmış ve dondurulmuştur.
v2.0 (Yeni Master Sürüm - 416K+ / 17 Kategori): Kodlar yalnızca sözdizimi… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/Turkish-Python-instruction.Turkish_corpus
Turkish Corpus 🇹🇷
Turkish Corpus is a large-scale cleaned Turkish text dataset created by collecting public Turkish corpora and extracting Turkish-language portions from multilingual datasets. The dataset is designed for Turkish Natural Language Processing research, language model pretraining, tokenizer training, embedding models, retrieval systems, and general Turkish language understanding tasks.
The main purpose of this dataset is to provide a practical, scalable, and… See the full description on the dataset page: https://huggingface.co/datasets/Ethosoft/Turkish_corpus.ru_turbo_alpaca
RuTurboAlpaca
Dataset of ChatGPT-generated instructions in Russian.
Code: rulm/self_instruct
Code is based on Stanford Alpaca and self-instruct.
29822 examples
Preliminary evaluation by an expert based on 400 samples:
83% of samples contain correct instructions
63% of samples have correct instructions and outputs
Crowdsouring-based evaluation on 3500 samples:
90% of samples contain correct instructions
68% of samples have correct instructions and outputs
Prompt template:… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/ru_turbo_alpaca.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/mrfg/turkish-court-decisions.mogan-turkish-web
Mogan Turkish Web
A large-scale Turkish web corpus derived from monthly Common Crawl snapshots
covering the period from January 2025 to June 2026. The corpus was
constructed by extracting Turkish-language content from raw Common Crawl
WARC/WET dumps, followed by language filtering, PII masking, and
near-duplicate removal.
📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum
Dataset Summary
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/moganai/mogan-turkish-web.turkish-corpus-100b
Turkish Corpus 100B (TC-100B)
Dataset Summary
The Turkish Corpus 100B (TC-100B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Turkish. Comprising approximately 105 Billion tokens (measured with Qwen/Llama3 tokenizer), it represents one of the largest open resources for Turkish LLM pretraining.
The dataset is engineered for a two-stage training pipeline:
Pretrain Subset (~103B Tokens): A diverse mix of… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/turkish-corpus-100b.ru_turbo_saiga
Saiga
Dataset of ChatGPT-generated chats in Russian.
Based on the Baize paper.
Code: link.
Prompt:
Идёт диалог между пользователем и ИИ ассистентом.
Пользователь и ассистент общаются на тему: {{seed}}
Реплики человека начинаются с [Пользователь], реплики ассистента начинаются с [Ассистент].
Пользователь задаёт вопросы на основе темы и предыдущих сообщений.
Пользователь обрывает беседу, когда у него не остается вопросов.
Ассистент даёт максимально полные, информативные, точные и… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/ru_turbo_saiga.turk-ictihat-kararlari-fulltext
Türk İçtihat Kararları - Tam Metin (UYAP e-Bedesten)
Adalet Bakanlığı UYAP Mevzuat ve İçtihat Programı (bedesten.adalet.gov.tr)
üzerinden derlenen Yargıtay kararları veri seti. Kavram bazlı arama
sorguları ile toplanmıştır.
Boyut
Kayıt: 9,899,589 benzersiz karar
Terim dağılımı: {"bosanma": 100052, "karar": 9894542, "nafaka": 66241, "tazminat": 100052}
Yıl aralığı: 1993-2026
Tam metni gelen karar: 39,301 (artımlı olarak dolduruluyor)
Şema… See the full description on the dataset page: https://huggingface.co/datasets/muhammedturan/turk-ictihat-kararlari-fulltext.temiz-mC4
Dataset Card for Temiz mC4
Temiz mC4 is the cleaned version of CulturaX corpus' Turkish split.
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
split
num instances
size
num of words
train
76.432.893
168GB
21.06B
Total
76.432.893
168GB
21.06B
This collection includes web text, crawled from internet for mC4 corpus. CulturaX is even refined version of mC4 with quality filtering and… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-mC4.TurMix
TurMix (https://arxiv.org/abs/2512.18834) is a Turkish pretraining corpus containing 168 billion tokens across 219 million documents (in the minhash subset). Rather than scraping the web again, TurMix combines five publicly available Turkish datasets, applies Turkish-specific quality filtering, and performs cross-dataset deduplication.
We train a 1.4B parameter language model through nanotron on 30 billion tokens to show that the matched subset of TurMix outperforms the… See the full description on the dataset page: https://huggingface.co/datasets/Alptekinege/TurMix.turkish-daily-dialogues-5k
Turkish Daily Dialogues 5K
Exactly 5,000 synthetic, multi-turn Turkish conversations covering ordinary daily-life situations. The corpus is designed as a small, auditable baseline for dialogue modelling, instruction-format experiments, augmentation research, and Turkish-language evaluation—not as a substitute for conversations written by real people.
Provenance in one sentence: the Turkish source scenario library was drafted with AI assistance specifically for this project… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/turkish-daily-dialogues-5k.temiz-OSCAR
Dataset Card for Temiz OSCAR
Temiz OSCAR is a corpora collection consisting of cleaned versions of original OSCAR corpora.
This collection is made up of four datasets: OSCAR-2019, OSCAR-2109, OSCAR-2201 and OSCAR-2301
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
Dataset
num instances
size
num of words
OSCAR-2019
3.671.430
7.7G
976M
OSCAR-2109
8.472.809
18G
2.22B
OSCAR-2201… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-OSCAR.Multi-turn_Long-context_Benchmark_for_LLMs
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
Arxiv: https://www.arxiv.org/abs/2507.13681
Huggingface: https://huggingface.co/papers/2507.13681
Introduction
LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios.
Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.TurkMorfBench
TurkMorfBench v3
465.241 madde · 14 kova · uydurma gövde kontrollü · teşhis raporlu
465,241 items · 14 buckets · wug-controlled · diagnostic reporting
eCloud Tech. · kod Apache-2.0 · veri CC BY 4.0
pip install turkmorfbench
turkmorfbench olc --model YOUR/MODEL
from datasets import load_dataset
d = load_dataset("ecloudtech/TurkMorfBench", "cekirdek") # 1.856 madde
d = load_dataset("ecloudtech/TurkMorfBench", "tam") # 465.241 madde
Ne ölçüyor
Türkçe… See the full description on the dataset page: https://huggingface.co/datasets/ecloudtech/TurkMorfBench.knowchat-multi-turn-dialogues
KnowChat: Multi-Turn Human-LLM Dialogues on Knowledge Tasks
KnowChat is a dataset of 705 multi-turn human-LLM conversations collected to validate the KnowSim user simulation framework. It pairs each conversation with pre/post knowledge assessments, self-reported survey ratings, and participant background information, enabling research on information calibration -- how well LLM assistants tailor responses to users with different knowledge levels.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/yjlee36/knowchat-multi-turn-dialogues.Turkish-SFT-Dataset-v1.0
Turkish-SFT-Dataset-v1.01
Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci
🔎 Özet
Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.turkish_parliamentary_data
Grand National Assembly Corpus of Türkiye (GNACT)
A comprehensive collection of Turkish parliamentary transcripts spanning over 100 years (1920–present), from 10 legislative bodies. Includes both Ottoman Turkish (1920–1928) and Modern Turkish (1928–present) texts.
Loading the dataset
from datasets import load_dataset
# Strategy 1: full session documents, all bodies (default)
ds = load_dataset("boun-tabilab/turkish_parliamentary_data", "full_sessions", split="train")
#… See the full description on the dataset page: https://huggingface.co/datasets/boun-tabilab/turkish_parliamentary_data.Turkce-Atlas-Instruct
Türkçe Atlas — Büyük Ölçekli Türkçe Instruct SFT Veri Kümesi
Türkçe Atlas, Türkçe komut takibi ve sohbet modeli eğitimi için hazırlanmış, konuşma biçiminde 336.146 örnek içeren bir denetimli ince ayar (Supervised Fine-Tuning, SFT) veri kümesidir. Her kayıt tek bir messages alanından oluşur ve sabit olarak system → user → assistant sırasındaki üç mesajı içerir.
60 kayıtlık düzenli örneklemde yeniden yazma, özetleme, soru-cevap, yapılandırılmış çıktı üretme, Türkçe dilbilgisi… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkce-Atlas-Instruct.turkce-sft-qa-3.7m
🇹🇷 Turkish SFT/QA — Birleştirilmiş ve Tekrarsız Veri Seti
3,723,264 örnek. 24 açık Türkçe SFT/QA veri setinin, satır düzeyinde
tekrar temizliği ve kalite kontrolünden geçirilmiş birleşimi. Her satır hangi veri
setinden geldiğini taşır.
English: A merged, row-level deduplicated and quality-filtered collection of
24 open Turkish SFT/QA datasets (3,723,264 examples). Every row carries
its source dataset, source URL and original license.
🙏 Teşekkür /… See the full description on the dataset page: https://huggingface.co/datasets/MercanAI/turkce-sft-qa-3.7m.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Alptekinege/turkish-court-decisions.qwen36-mtp-turbo-kv-analysis
Qwen3.6 MTP Turbo KV Runtime Analysis
This repository is a curated analysis artifact for local Qwen3.6-35B-A3B MTP GGUF inference experiments on Windows CUDA. It compares clean MTP llama.cpp, QuinsZouls llama-next TurboQuant, and the completed subset of Atomic TurboQuant runs under a fixed 64k context, MoE CPU offload, and Unsloth-aligned sampling settings.
The raw benchmark runs included incomplete and capability-incompatible rows. This repo keeps only completed, comparable… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwen36-mtp-turbo-kv-analysis.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Gyrevortex/turkish-court-decisions.turkish-llm-injection
🇹🇷 AltaySec Turkish LLM Prompt Injection Dataset (v0.2)
Türkiye'nin ilk Türkçe-öncelikli, kategorize edilmiş LLM prompt injection veri seti — genişletilmiş sürüm.
📌 TL;DR
300 elle/üretim-destekli hazırlanmış Türkçe prompt injection payload'u, 12 saldırı kategorisi × 25, OWASP LLM Top 10 (2025) ile eşlenmiş. v0.1'in 120 çekirdek payload'una, AltayDuel arenasındaki bulgular ışığında üretilip düşmanca kalite/dedup denetiminden geçirilmiş 180 yeni payload eklendi.… See the full description on the dataset page: https://huggingface.co/datasets/AltaySec/turkish-llm-injection.AkademikDerlem
Dataset Card for AkademikDerlem
AkademikDerlem is a scientific text corpus for Turkish, gathered from misc academical publication websites.
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
This collection is made up of five datasets: Articles, Academic-Abstracts, Medical-Articles, Medical-Abstracts, and Bilkent-Writings. The Bilkent-Writings dataset comes from creative writings produced in the… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/AkademikDerlem.Flutter-Code-with-Questions-Dataset-Turkish
Flutter Code with Questions Dataset (Turkish)
📦 Dataset Name: flutter_code_with_questions
Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir.
📁 Dataset Format
Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.Havadis
Dataset Card for Havadis
Havadis is a high quality and large Turkish news corpus, indeed the largest Turkish news corpus ever.
This corpus is scraped from online news sebsites and includes text from popular newspapers such as
CNN Türk
Habertürk
Hürriyet
Millyet
NTV
Posta
Sabah
Star
Sözcü
Takvim
. The instances are first crawled from the corresponding websites, then went throught an extensive cleaning process. We eliminated instances that are too short, too repetetive… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Havadis.
