datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish_political_position_benchmark
Turkish Political Position Benchmark
The Turkish Political Position Benchmark measures how language models respond to normative statements about Turkish politics. It reports ideological dimension scores and response similarity to documented political-party reference profiles.
The benchmark does not claim that a model belongs to a party, has a voting intention, or possesses political beliefs. A party similarity score only means that the model produced a similar pattern of answers… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/turkish_political_position_benchmark.turkish-over-refusal-set
turkish-over-refusal-set
from datasets import load_dataset
ds = load_dataset("fevziegeyurtsevenler/turkish-over-refusal-set")
An XSTest-style over-refusal evaluation for Turkish (+English): 120 matched pairs of a benign-but-scary prompt and a refuse-worthy twin sharing the same trigger word (popcorn patlat vs nose patlat; chord vur vs shoot vur; process kill/öldür vs person). 480 prompts, 10 categories.
Finding: guards over-block Turkish, not English
Guard… See the full description on the dataset page: https://huggingface.co/datasets/fevziegeyurtsevenler/turkish-over-refusal-set.turkish-tool-calling
Türkçe Tool-Calling Veri Seti
56.247 kayıt. xLAM/APIGen 60k ve NVIDIA When2Call'dan türetilmiş,
üç davranış sınıfı içeren Türkçe function-calling veri seti.
from datasets import load_dataset
ds = load_dataset("bilalabic/turkish-tool-calling") # mesaj listesi
ds = load_dataset("bilalabic/turkish-tool-calling", "table") # düz tablo
ds = load_dataset("bilalabic/turkish-tool-calling", "sharegpt") # ShareGPT
İçerik
Kayıt
56.247… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/turkish-tool-calling.turkish-legal-statutory-hallucination-benchmark
Citation
If you use this dataset, please cite the accompanying paper:
@inproceedings{erdoganyilmaz2026statutoryhallucinations,
title = {Measuring Statutory Citation Hallucinations of LLMs in Turkish Law: A Multi-Agent Based Novel Benchmark Dataset and Multi-Dimensional Evaluation Framework},
author = {Cihan Erdoğanyılmaz and Ali Yasir Naç and Gamze Çoskuner},
booktitle = {2026 34th Signal Processing and Communications Applications Conference (SIU)},
year =… See the full description on the dataset page: https://huggingface.co/datasets/LawChatAI/turkish-legal-statutory-hallucination-benchmark.Turkish-Chat_GPT-4O
Quardo/Turkish-Chat_GPT-4O
Description
This is a simple dataset generated by OpenAI's GPT-4O (gpt-4o-2024-08-06). The dataset includes various entries created and evaluated by the AI model, providing a unique collection of Turkish chat data for analysis and research.
Warning
Please note that this dataset may contain errors or inconsistencies as it is fully generated by an AI model. It is highly recommended to check and edit the data before usage, as AI can… See the full description on the dataset page: https://huggingface.co/datasets/Quardo/Turkish-Chat_GPT-4O.turkish-seo-reasoning-benchmark-results
Turkish SEO Reasoning Benchmark Results
Bu dataset, Turkish SEO Reasoning benchmark'ının altı farklı model/checkpoint üzerinde çalıştırılmış ham tahminlerini, metriklerini ve tekrar üretim manifestlerini içerir.
Fine-tuned model: berkbirkan/gemma-3-1b-turkish-seo-reasoning-lora
Sonuç
Fine-tuned Gemma 3 1B modeli 22,23 skorla ilk sırada yer aldı. Aynı base model 11,96 skor elde etti.
Mutlak artış: +10,28 puan
Göreli artış: %85,97
Fine-tuned model hata sayısı:… See the full description on the dataset page: https://huggingface.co/datasets/berkbirkan/turkish-seo-reasoning-benchmark-results.TurkishMMLU-Kirlilik
TurkishMMLU Kirlilik Listesi — Türkçe web'de aynen bulunan test soruları
TurkishMMLU test kümesindeki 900 sorunun 113'ü (%12,6) 527 milyon kelimelik
bir Türkçe web külliyatında birebir geçmektedir. Bu depo o soruların listesini,
bulunma yöntemini ve yeniden üretim betiğini içerir; amaç, Türkçe web metniyle
eğitilen modellerin TurkishMMLU puanlarını kirlilik dışı bırakılmış bir alt kümede
de raporlayabilmesidir.
English summary at the end.
Neden önemli
Bir kıyas… See the full description on the dataset page: https://huggingface.co/datasets/ecloudtech/TurkishMMLU-Kirlilik.turkish-wikipedia-dataset-clean
Turkish Wikipedia Dataset
A cleaned and structured Turkish Wikipedia dataset designed for Turkish language model pretraining, continued pretraining, research, and NLP experiments.
The dataset consists of articles collected from the Turkish Wikipedia (tr.wikipedia.org) and processed into a machine-readable format while preserving important source metadata.
Dataset Summary
Language: Turkish (tr)
Source: Turkish Wikipedia
Domain: General knowledge / encyclopedia… See the full description on the dataset page: https://huggingface.co/datasets/kaan39/turkish-wikipedia-dataset-clean.instruction-turkish-poems
Turkish poems for fine-tuning LLMs with instructions.
Instructions created with Google's Gemini-Pro.
For a dataset that has variety of instructions check: beratcmn/rephrased-instruction-turkish-poems
Base dataset:
beratcmn/turkish-poems-cleaned
rephrased-instruction-turkish-poems
This is a rephrased version of my previous dataset beratcmn/instruction-turkish-poems. I used the same instructions but I rephrased them to be more clear and understandable also added more variety to the format.
devim-grounded-turkish
DEVİM Grounded Turkish
This is a methodology and public-evidence repository, not a release of the underlying rights-restricted Turkish corpus.
DEVİM's grounded-data program was created to reduce shortcut learning and weak transfer by linking supervision to source evidence, preserving provenance, separating training material from sequestered evaluation material, and explicitly testing abstention when an answer is not supported.
Verified source frame
Authorized… See the full description on the dataset page: https://huggingface.co/datasets/bazobehram/devim-grounded-turkish.turkish-question-departmentturkish-poems-cleanedHTML tags removed and overall cleaned version of okg/turkish-poems. Original: https://huggingface.co/datasets/okg/turkish-poems
Data için teşekkür ederim okg <3
nocisnn-turkish-logic-traces
NociSNN Turkish Mathematical Logic Traces
This dataset records structured Turkish mathematical reasoning traces generated by
ogulcanaydogan/Turkish-LLM-7B-Instruct-GGUF:Q4_K_M in a continuously running
NociSNN evaluation loop.
Data Generation
Problems are generated from deterministic Turkish templates covering arithmetic,
percentage, proportion and single-variable linear equations. Ground-truth
answers are computed by the task generator rather than the language… See the full description on the dataset page: https://huggingface.co/datasets/Lightcap/nocisnn-turkish-logic-traces.Turkish_HellaSwag10_WrongsGPT-5.4-Turkish-ReasoningTurkish_HellaswagHellaSwag_Turkish_Choices_0HellaSwag_Turkish_Choices_2Turkish_HellaSwag10_Shuffle_FirstHellaSwag_Turkish_Choices_1HellaSwag_Turkish_Choices_3Turkish_HellaSwag10_Shuffle_SecondTurkish_HellaSwag10_Shuffle_ThirdTurkishAcademic-Datasetdergipark-turkish-cleaned
