datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkishfineweb2-cleaned
TurkishFineweb2-Cleaned
A Turkish web corpus derived from the Turkish (tur_Latn) subset of
FineWeb-2, augmented with an additional quality-classification layer
and a near-duplicate removal pass.
📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum
Source
FineWeb-2 is a
large-scale, multilingual web corpus built from Common Crawl. This dataset
covers the Turkish (tur_Latn) portion of FineWeb-2, spanning the… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkishfineweb2-cleaned.Turkish_corpus
Turkish Corpus 🇹🇷
Turkish Corpus is a large-scale cleaned Turkish text dataset created by collecting public Turkish corpora and extracting Turkish-language portions from multilingual datasets. The dataset is designed for Turkish Natural Language Processing research, language model pretraining, tokenizer training, embedding models, retrieval systems, and general Turkish language understanding tasks.
The main purpose of this dataset is to provide a practical, scalable, and… See the full description on the dataset page: https://huggingface.co/datasets/Ethosoft/Turkish_corpus.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/mrfg/turkish-court-decisions.turkish_parliamentary_data
Grand National Assembly Corpus of Türkiye (GNACT)
A comprehensive collection of Turkish parliamentary transcripts spanning over 100 years (1920–present), from 10 legislative bodies. Includes both Ottoman Turkish (1920–1928) and Modern Turkish (1928–present) texts.
Loading the dataset
from datasets import load_dataset
# Strategy 1: full session documents, all bodies (default)
ds = load_dataset("boun-tabilab/turkish_parliamentary_data", "full_sessions", split="train")
#… See the full description on the dataset page: https://huggingface.co/datasets/boun-tabilab/turkish_parliamentary_data.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Alptekinege/turkish-court-decisions.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Gyrevortex/turkish-court-decisions.turkish-court-decisions-duplicate
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-court-decisions-duplicate.fineweb-2-turkish-categorized-long
altaidevorg/fineweb-2-turkish-categorized long filtered Turkish texts
Source: altaidevorg/fineweb-2-turkish-categorized (config: default).
The script streamed 10,000,000 raw source rows before stopping. Categories
ads, adult content, sports, tabloid were rejected before length and quality filtering.
Retained rows contain 3,000–16,500 characters
and passed the iteration-5 Turkish language,
repetition, glue-word, punctuation, SEO, and soft information-density filters.
Selected… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/fineweb-2-turkish-categorized-long.Turkish-Recipe-Corpus
Turkish Recipe Corpus (TRC-30K)
Türkçe'nin en kapsamlı açık kaynaklı tarif veri seti.74,768 Türkçe tariften oluşan, 3 farklı NLP görevine hazır yapılandırılmış corpus.
Ethosoft Research · huggingface.co/Ethosoft · ethosoft.org
Dataset Özeti
Değer
Toplam tarif
74,768
Dil
Türkçe (tr)
Lisans
CC BY 4.0
Konfigürasyonlar
3 (structured, instruction, ingredient2recipe)
Split
Train / Validation / Test (80 / 10 / 10)
Ortalama adım sayısı
5.23 /… See the full description on the dataset page: https://huggingface.co/datasets/Ethosoft/Turkish-Recipe-Corpus.turkish-brand-bias-evaluations
Turkish Brand Bias Evaluations / Türkçe Marka Yanlılığı Değerlendirmeleri
Furkan Karlı tarafından Türkçe ürün ve hizmet önerilerindeki marka görünürlüğünü
incelemek amacıyla oluşturulmuş LLM değerlendirme veri setidir.
An LLM evaluation dataset curated by Furkan Karlı to study brand visibility in
Turkish product and service recommendations.
Veri seti özeti
300 tamamlanmış ve judge edilmiş yanıt
Domainler: VPN (150) ve kozmetik (150)
Koşullar: web araması kapalı… See the full description on the dataset page: https://huggingface.co/datasets/furkankarli/turkish-brand-bias-evaluations.turkish-tool-calling
Türkçe Tool-Calling Veri Seti
56.247 kayıt. xLAM/APIGen 60k ve NVIDIA When2Call'dan türetilmiş,
üç davranış sınıfı içeren Türkçe function-calling veri seti.
from datasets import load_dataset
ds = load_dataset("bilalabic/turkish-tool-calling") # mesaj listesi
ds = load_dataset("bilalabic/turkish-tool-calling", "table") # düz tablo
ds = load_dataset("bilalabic/turkish-tool-calling", "sharegpt") # ShareGPT
İçerik
Kayıt
56.247… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/turkish-tool-calling.old-nogay-turkish-ocr-corpus
Old Nogay Turkish OCR Corpus
This is a small OCR-derived corpus of historical Nogay Turkish / Turkic textual material, collected, classified, extracted, and packaged by Fatih Burak Karagöz / CDLI.ai for exploratory NLP, historical corpus work, and OCR-quality analysis.
We are proud to release this as a contribution to Turkish NLP, Turkic-language NLP, historical NLP, and Turkish studies. The goal is not to pretend that a small OCR corpus is a polished benchmark. The goal is more… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/old-nogay-turkish-ocr-corpus.Turkish-Chat_GPT-4O
Quardo/Turkish-Chat_GPT-4O
Description
This is a simple dataset generated by OpenAI's GPT-4O (gpt-4o-2024-08-06). The dataset includes various entries created and evaluated by the AI model, providing a unique collection of Turkish chat data for analysis and research.
Warning
Please note that this dataset may contain errors or inconsistencies as it is fully generated by an AI model. It is highly recommended to check and edit the data before usage, as AI can… See the full description on the dataset page: https://huggingface.co/datasets/Quardo/Turkish-Chat_GPT-4O.turkish-seo-reasoning-benchmark-results
Turkish SEO Reasoning Benchmark Results
Bu dataset, Turkish SEO Reasoning benchmark'ının altı farklı model/checkpoint üzerinde çalıştırılmış ham tahminlerini, metriklerini ve tekrar üretim manifestlerini içerir.
Fine-tuned model: berkbirkan/gemma-3-1b-turkish-seo-reasoning-lora
Sonuç
Fine-tuned Gemma 3 1B modeli 22,23 skorla ilk sırada yer aldı. Aynı base model 11,96 skor elde etti.
Mutlak artış: +10,28 puan
Göreli artış: %85,97
Fine-tuned model hata sayısı:… See the full description on the dataset page: https://huggingface.co/datasets/berkbirkan/turkish-seo-reasoning-benchmark-results.Instruction-280K-Turkish
Dataset Card for Instruction-280K-Turkish
Language: Turkish
Dataset Description
This repository contains a dataset for Turkish version of the Deepseek 1.5B model. The translation was performed using the Google translation model to ensure high-quality, accurate translation.
Dataset Details
Size: ≈280K
Translation tool: Google Translate
Data format: Prompt, Response
turkish-poemsTurkish poems scraped from antoloji.com. Features consists of id, poet name, poem rating and the poem.
Turkish-Lyric-Intelligence-v2
Turkish Lyric Intelligence v2
Turkish Lyric Intelligence v2 is a derived, research-oriented dataset for
building controllable Turkish lyric generation and lyric-analysis systems. It
adds section structure, orthographic syllable counts, rhyme-ending candidates,
repetition signals, and review queues to the source
Genius Turkish Dataset.
This release is intended as an intermediate data layer. Automated rhyme,
prosody, and task labels are candidate annotations, not expert literary… See the full description on the dataset page: https://huggingface.co/datasets/mustafakemal0146/Turkish-Lyric-Intelligence-v2.mogan-turkish-web-long
moganai/mogan-turkish-web long filtered Turkish texts
Source: moganai/mogan-turkish-web (config: default, revision: d773a0efd1b7daf72c7909c83dbd385e7e3564d7).
Rows contain 4,000–16,000 characters and passed the
iteration-5 Turkish language, repetition, glue-word, punctuation, SEO, and soft
information-density filters. Selected rows: 2,805,010.
Generated by process_hf_dataset.py. See summary.json for counts and thresholds.
rephrased-instruction-turkish-poems
This is a rephrased version of my previous dataset beratcmn/instruction-turkish-poems. I used the same instructions but I rephrased them to be more clear and understandable also added more variety to the format.
turkish-math-rlvr
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti barandinho tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: barandinho/turkish-math-rlvr
🔗 Derleyen Platform: VeriPazarı
Türkçe Matematiksel Akıl Yürütme (RLVR Eğitim Veri Seti)
Bu Veri Seti Nedir?
Bu veri seti, zayıf bir modelin başarı oranına… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/turkish-math-rlvr.Turkish-Legislation-DPO
Turkish-Legislation-DPO Dataset
Turkish-Legislation-DPO is a large-scale Direct Preference Optimization (DPO) training dataset specifically designed to align Turkish language models with expertise in the fields of law and regulation.
Dataset Summary
The Turkish-Legislation-DPO dataset contains 23,596 carefully selected preference pairs, all generated by google/gemma-2-2b-it model and improved through systematic quality assessment protocols. Each example consists of… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-Legislation-DPO.turkish-number-words-1m
Turkish Number Words 1M v2
0 ile 999.999 arasındaki her tamsayının Türkçe yazıyla karşılığı.
Doğrulanmış boyut
Train: 980,000
Validation: 10,000
Test: 10,000
Toplam: 1,000,000
Ana görev sütunları: id, number, words
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance, generator_version, generator_sha256… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-number-words-1m.nocisnn-turkish-logic-traces
NociSNN Turkish Mathematical Logic Traces
This dataset records structured Turkish mathematical reasoning traces generated by
ogulcanaydogan/Turkish-LLM-7B-Instruct-GGUF:Q4_K_M in a continuously running
NociSNN evaluation loop.
Data Generation
Problems are generated from deterministic Turkish templates covering arithmetic,
percentage, proportion and single-variable linear equations. Ground-truth
answers are computed by the task generator rather than the language… See the full description on the dataset page: https://huggingface.co/datasets/Lightcap/nocisnn-turkish-logic-traces.Turkish-STEM-DPO-Dataset
Turkish STEM DPO Dataset
Dataset Summary
The Turkish STEM DPO (Direct Preference Optimization) dataset is a comprehensive synthetic resource containing 16,177 high-quality preference pairs designed to enhance the reasoning capabilities of Turkish language models in mathematics, physics, and programming.
The dataset leverages a preference-based learning approach: each instance pairs a carefully crafted, expert-level solution with a deliberately flawed or incomplete… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-STEM-DPO-Dataset.Turkish-STEM-DPO-Dataset
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti yusufbaykaloglu tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: yusufbaykaloglu/Turkish-STEM-DPO-Dataset
🔗 Derleyen Platform: VeriPazarı
Türkçe STEM DPO Veri Seti (Turkish STEM DPO Dataset)
Veri Seti Özeti
Turkish STEM DPO (Doğrudan Tercih… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/Turkish-STEM-DPO-Dataset.
