datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Turkish-Python-instruction
🚀 DİKKAT VERİ SETİ GÜNCELLENME SÜRECİNE ALINMIŞTIR LÜTFEN AÇIKLAMAYI OKUYUNUZ. Turkish Python & System Engineering Dataset (BYSISMO v2.0)
25 Kategorilik Büyük Türkçe Python & Sistem Mühendisliği Havuzu
📢 SÜRÜM & DOĞRULAMA DURUMU (VERSION ROADMAP)
v1.0 (Eski Arşiv - 289K / 8 Kategori): Yüksek kalite standartlarımız gereği yeniden yapılandırmaya alınmış ve dondurulmuştur.
v2.0 (Yeni Master Sürüm - 416K+ / 17 Kategori): Kodlar yalnızca sözdizimi… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/Turkish-Python-instruction.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/mrfg/turkish-court-decisions.turkish-corpus-100b
Turkish Corpus 100B (TC-100B)
Dataset Summary
The Turkish Corpus 100B (TC-100B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Turkish. Comprising approximately 105 Billion tokens (measured with Qwen/Llama3 tokenizer), it represents one of the largest open resources for Turkish LLM pretraining.
The dataset is engineered for a two-stage training pipeline:
Pretrain Subset (~103B Tokens): A diverse mix of… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/turkish-corpus-100b.Multi-Turn-Insurance-Underwriting
Dataset Card for Multi-Turn-Insurance-Underwriting
Dataset Summary
This dataset includes sample traces and associated metadata from multi-turn interactions between a commercial underwriter and AI assistant. We built the system in langgraph with model context protocol and ReAct agents. In each sample, the underwriter has a specific task to solve related to a recent application for insurance by a small business. We created a diverse sample dataset covering 6 distinct types… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Multi-Turn-Insurance-Underwriting.STRIDE-QA-Dataset-Mini
STRIDE-QA-Dataset-Mini
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
⚠️ Note: STRIDE-QA-Dataset-Mini is provided as a preliminary version and does not fully match the format of the… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset-Mini.Turkish-SFT-Dataset-v1.0
Turkish-SFT-Dataset-v1.01
Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci
🔎 Özet
Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.knowchat-multi-turn-dialogues
KnowChat: Multi-Turn Human-LLM Dialogues on Knowledge Tasks
KnowChat is a dataset of 705 multi-turn human-LLM conversations collected to validate the KnowSim user simulation framework. It pairs each conversation with pre/post knowledge assessments, self-reported survey ratings, and participant background information, enabling research on information calibration -- how well LLM assistants tailor responses to users with different knowledge levels.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/yjlee36/knowchat-multi-turn-dialogues.Multi-turn_Long-context_Benchmark_for_LLMs
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
Arxiv: https://www.arxiv.org/abs/2507.13681
Huggingface: https://huggingface.co/papers/2507.13681
Introduction
LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios.
Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.TurkishMMLU
TurkishMMLU
This repository contains Code and Data Analysis of TurkishMMLU for ACL 24 SIGTURK Workshop. TurkishMMLU is a multiple-choice dataset for Turkish Natural Language Processing (NLP) community based on Turkish Highschool Curricula for nine subjects.
To access this dataset please send an email to:
arda.yueksel@tum.de or akoksal@cis.lmu.de.
Abstract
Multiple choice question answering tasks evaluate the reasoning, comprehension, and mathematical abilities of… See the full description on the dataset page: https://huggingface.co/datasets/AYueksel/TurkishMMLU.turkish_parliamentary_data
Grand National Assembly Corpus of Türkiye (GNACT)
A comprehensive collection of Turkish parliamentary transcripts spanning over 100 years (1920–present), from 10 legislative bodies. Includes both Ottoman Turkish (1920–1928) and Modern Turkish (1928–present) texts.
Loading the dataset
from datasets import load_dataset
# Strategy 1: full session documents, all bodies (default)
ds = load_dataset("boun-tabilab/turkish_parliamentary_data", "full_sessions", split="train")
#… See the full description on the dataset page: https://huggingface.co/datasets/boun-tabilab/turkish_parliamentary_data.turkce-sft-qa-3.7m
🇹🇷 Turkish SFT/QA — Birleştirilmiş ve Tekrarsız Veri Seti
3,723,264 örnek. 24 açık Türkçe SFT/QA veri setinin, satır düzeyinde
tekrar temizliği ve kalite kontrolünden geçirilmiş birleşimi. Her satır hangi veri
setinden geldiğini taşır.
English: A merged, row-level deduplicated and quality-filtered collection of
24 open Turkish SFT/QA datasets (3,723,264 examples). Every row carries
its source dataset, source URL and original license.
🙏 Teşekkür /… See the full description on the dataset page: https://huggingface.co/datasets/MercanAI/turkce-sft-qa-3.7m.Turkce-Atlas-Instruct
Türkçe Atlas — Büyük Ölçekli Türkçe Instruct SFT Veri Kümesi
Türkçe Atlas, Türkçe komut takibi ve sohbet modeli eğitimi için hazırlanmış, konuşma biçiminde 336.146 örnek içeren bir denetimli ince ayar (Supervised Fine-Tuning, SFT) veri kümesidir. Her kayıt tek bir messages alanından oluşur ve sabit olarak system → user → assistant sırasındaki üç mesajı içerir.
60 kayıtlık düzenli örneklemde yeniden yazma, özetleme, soru-cevap, yapılandırılmış çıktı üretme, Türkçe dilbilgisi… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkce-Atlas-Instruct.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Alptekinege/turkish-court-decisions.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Gyrevortex/turkish-court-decisions.Flutter-Code-with-Questions-Dataset-Turkish
Flutter Code with Questions Dataset (Turkish)
📦 Dataset Name: flutter_code_with_questions
Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir.
📁 Dataset Format
Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.Vietnamese-Multi-turn-Chat-AlpacaInstrucTurca
InstrucTurca v1.0.0 is a diverse synthetic instruction tuning dataset crafted for instruction-tuning Turkish LLMs. The data is compiled data various English datasets and sources, such as code instructions, poems, summarized texts, medical texts, and more.
Dataset content
BI55/MedText
checkai/instruction-poems
garage-bAInd/Open-Platypus
Locutusque/ColumnedChatCombined
nampdn-ai/tiny-codes
Open-Orca/OpenOrca
pubmed_qa
TIGER-Lab/MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/InstrucTurca.turkish-court-decisions-duplicate
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-court-decisions-duplicate.Open-MM-RL
Dataset Summary
Open-MM-RL is a multimodal STEM reasoning dataset covering Physics, Mathematics, Biology, and Chemistry. It is designed for problems that require models to interpret visual information and combine it with step-by-step analytical reasoning.
Explore the full Open-MM-RL dataset (3,000 tasks coming soon): https://go.turing.com/open-mm-rl
Compared with existing multimodal reasoning benchmarks, Open-MM-RL broadens the evaluation setting beyond standard single-image… See the full description on the dataset page: https://huggingface.co/datasets/TuringEnterprises/Open-MM-RL.turkish-law-corpus
⚖️ Turkish Law — 106 Kanun Korpusu & Soru-Cevap106 Statutes Corpus & QA
🇹🇷 Türk hukukunun en çok kullanılan 106 kanunu, madde madde temizlenmiş 16.001 metin parçası ve bu maddelere dayalı 5.011 Türkçe soru-cevap çifti. Tamamı resmî kaynaktan (mevzuat.gov.tr), RAG ve yapay zekâ uygulamaları için hazır.
🇬🇧 The 106 most widely used Turkish statutes as 16,001 clean, article-level text chunks, plus 5,011 Turkish question-answer pairs grounded in those articles. All from the… See the full description on the dataset page: https://huggingface.co/datasets/CtnkyaABC/turkish-law-corpus.turkish-extractive-qa-1.5m
Turkish Extractive QA 1.5M v2
Cevap metni ve başlangıç konumu doğrulanabilir Türkçe çıkarımsal soru-cevap kayıtları.
Doğrulanmış boyut
Train: 1,470,000
Validation: 15,000
Test: 15,000
Toplam: 1,500,000
Ana görev sütunları: id, context, question, answer, answer_start, question_type
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-extractive-qa-1.5m.turkish-medicine-law
turkish-medicine-law
Bu veri seti, Türkçe tıp ve sağlık hukuku alanında hazırlanmıştır. Türkçe hukuk alanında genel amaçlı birkaç kaynak bulunuyor, ama tıp hukuku özelinde hazırlanmış bir veri seti şimdiye kadar yoktu. Bu proje o boşluğu doldurmayı amaçlıyor.
Veri setindeki örnekler hukukçular, bilirkişiler ve sağlık kuruluşlarının hukuk birimleri için hazırlandı. Hastaya veya hekime doğrudan hukuki görüş sunmak amacıyla kullanılmak üzere tasarlanmadı. Buradaki çıktılar bir ön… See the full description on the dataset page: https://huggingface.co/datasets/tunahanf/turkish-medicine-law.Bilge-Turkish-CoT-50K
Bilge: Turkish Chain-of-Thought Dataset (50K)
50,000 örneklik Türkçe Chain-of-Thought (CoT) reasoning fine-tuning veri seti.
Bilge, Türkçe büyük dil modellerinin adım adım düşünme (reasoning) kapasitesini
geliştirmek amacıyla hazırlanmış bir Chain-of-Thought veri setidir.
Veri setindeki her örnek, modelin önce <think> blokları içinde görünür bir
muhakeme süreci yürütmesini, ardından kullanıcıya yapılandırılmış ve detaylı
bir cevap vermesini öğretmek üzere tasarlanmıştır.
Bu… See the full description on the dataset page: https://huggingface.co/datasets/bugrabilge/Bilge-Turkish-CoT-50K.finbenchv2-fbv1-stripped-fi-htThis is a subset of tasks from FIN-bench with prompts stripped from the inputs to be used within the FIN-bench-v2 benchmark suite as described in FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models.
Github: https://github.com/LumiOpen/lm-evaluation-harness
UPDATE - 2025-07-24:The format of similarities_abstraction_zero_shot has been changed to only include the word pair.sentence_ambiguity_zero_shot has been removed as it was broken and should not… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/finbenchv2-fbv1-stripped-fi-ht.turkish-competition-authority-decisions
Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026
The complete published decision history of the Turkish Competition Authority
(Rekabet Kurumu) — every Competition Board decision the regulator has made public,
in full text, with derived structural metadata.
10,367 decisions · 113,297 pages · 323 million characters · 29 years
Every decision carries its outcome, the articles of Law 4054 it turns on, the
panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/emirms/turkish-competition-authority-decisions.Rubric-Graded-Reasoning
Rubrics-Graded Reasoning — Computer Science, Data Science, Chemistry
A multi-domain reasoning dataset built to improve frontier models by revealing their failures and turning expert grading into training signal.
The dataset pairs self-contained tasks with weighted rubrics across three domains — Computer Science, Data Science, and Chemistry — turning expert evaluation into training signals that boost frontier-model reasoning.
Explore the full Rubric-based reasoning data pack:… See the full description on the dataset page: https://huggingface.co/datasets/TuringEnterprises/Rubric-Graded-Reasoning.Open-RL
Open-RL
Dataset Summary
This dataset contains self-contained, verifiable, and unambiguous STEM reasoning problems across Physics, Mathematics, Biology, and Chemistry.
Each problem:
Requires multi-step reasoning
Involves symbolic manipulation and/or numerical computation
Has a deterministic, objectively verifiable final answer
The problems were evaluated against contemporary large language models. Observed pass rates indicate that the tasks are non-trivial yet… See the full description on the dataset page: https://huggingface.co/datasets/TuringEnterprises/Open-RL.turkish-medical-rag
🩺 Turkish Medical RAG
Hierarchical Parent–Child Retrieval-Augmented Generation for Turkish Medical Documents
📌 Proje Hakkında
Bu proje, Türkçe tıbbi dokümanlar üzerinde çalışan uçtan uca bir
Retrieval-Augmented Generation (RAG) sistemi geliştirmek amacıyla hazırlanmıştır.
Sistem bir kullanıcı sorusu aldığında önce doküman koleksiyonundaki küçük ve
anlamsal olarak odaklı parçalar (child chunks)… See the full description on the dataset page: https://huggingface.co/datasets/sedayzc/turkish-medical-rag.TurtleBench1.5k
Overview
TurtleBench is a novel evaluation benchmark designed to assess the reasoning capabilities of large language models (LLMs) using yes/no puzzles (commonly known as "Turtle Soup puzzles"). This dataset is constructed based on user guesses collected from our online Turtle Soup Puzzle platform, providing a dynamic and interactive means of evaluation. Unlike traditional static evaluation benchmarks, TurtleBench focuses on testing models in interactive settings to better capture… See the full description on the dataset page: https://huggingface.co/datasets/Duguce/TurtleBench1.5k.Turing-Open-Reasoning
Computational STEM QA Dataset
Dataset Summary
This dataset contains computationally intensive, self-contained, and unambiguous STEM reasoning problems across Physics, Mathematics, Biology, and Chemistry.
Problems require multi-step reasoning, symbolic manipulation, numerical accuracy, or simulation-based verification. These tasks expose failure modes in state-of-the-art LLMs, making this dataset a strong benchmark for evaluating deep reasoning.
Each example includes:… See the full description on the dataset page: https://huggingface.co/datasets/TuringEnterprises/Turing-Open-Reasoning.
