datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Flutter-Code-with-Questions-Dataset-Turkish
Flutter Code with Questions Dataset (Turkish)
📦 Dataset Name: flutter_code_with_questions
Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir.
📁 Dataset Format
Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.turkish-law-corpus
⚖️ Turkish Law — 106 Kanun Korpusu & Soru-Cevap106 Statutes Corpus & QA
🇹🇷 Türk hukukunun en çok kullanılan 106 kanunu, madde madde temizlenmiş 16.001 metin parçası ve bu maddelere dayalı 5.011 Türkçe soru-cevap çifti. Tamamı resmî kaynaktan (mevzuat.gov.tr), RAG ve yapay zekâ uygulamaları için hazır.
🇬🇧 The 106 most widely used Turkish statutes as 16,001 clean, article-level text chunks, plus 5,011 Turkish question-answer pairs grounded in those articles. All from the… See the full description on the dataset page: https://huggingface.co/datasets/CtnkyaABC/turkish-law-corpus.TARA_Turkish_LLM_Benchmark
TARA: Turkish Advanced Reasoning Assessment Veri Seti
*Img Credit: Open AI ChatGPT
**English version is given below.**
Evaluation Notebook / Değerlendirme Not Defteri
Dataset Summary
TARA (Turkish Advanced Reasoning Assessment), Türkçe dilindeki Büyük Dil Modellerinin (LLM'ler) gelişmiş akıl yürütme yeteneklerini çoklu alanlarda ölçmek için tasarlanmış, zorluk derecesine göre sınıflandırılmış bir benchmark veri setidir. Bu veri seti, LLM'lerin sadece bilgi… See the full description on the dataset page: https://huggingface.co/datasets/emre/TARA_Turkish_LLM_Benchmark.turkish-legal-rag
Turkish Legal RAG Corpus — Türk Hukuku için Açık RAG Datasetı
Tek cümle: 25 önemli Türk kanununun (mevzuat.gov.tr kaynaklı, madde bazlı temiz chunk'lar) + 290 manuel doğrulanmış soru-cevap altın benchmark'ının olduğu açık kaynak Türkçe hukuk RAG datasetı.
🇹🇷 Türkçe Özet — Bu dataset, Türkçe hukuk uygulamaları için sıfırdan üretilmiş açık ve denetlenebilir bir RAG corpus'udur. mevzuat.gov.tr üzerinden alınan 25 ana kanunun madde madde temizlenmiş, chunk'lanmış sürümünü (6.350… See the full description on the dataset page: https://huggingface.co/datasets/mtntasci/turkish-legal-rag.easy_turkish_math_reasoning
Easy Turkish Math Reasoning
Dataset Summary
The Easy Turkish Math Reasoning dataset is the first phase of a multi-stage curriculum learning pipeline designed to enhance the reasoning abilities of compact language models. This dataset focuses on elementary-level arithmetic and logic problems in Turkish, serving as a warm-up stage for supervised fine-tuning (SFT).
Use Case
Primarily used for:
Bootstrapping reasoning ability in Turkish for compact LLMs.
Phase 1… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/easy_turkish_math_reasoning.medium_turkish_math_reasoning
Dataset Summary
The Medium Turkish Math Reasoning dataset is Phase 2 of a curriculum learning pipeline to teach compact models multi-step reasoning in Turkish. It includes moderately difficult math problems involving multiple reasoning steps, such as two-part arithmetic, comparisons, and logical reasoning.
Use Case
This dataset is ideal for:
Continuing SFT after foundational training with simpler problems.
Bridging the gap between basic arithmetic and complex GSM8K-style… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/medium_turkish_math_reasoning.turkish_llm_finetune_dataset_4_topics
Turkish LLM Finetune Dataset - 4 Topics
This dataset is designed to fine-tune the T3 AI Turkish LLM. It was created by Barathan Aslan, Ömer Faruk Çelik, and Batuhan Kalem for the T3 AI Hackathon. The dataset focuses on four distinct topics: Agriculture, Sustainability, Turkish Education Sytem, and Turkish Law System.
Contributors
Barathan Aslan (https://huggingface.co/barathanasln)
Batuhan Kalem(https://huggingface.co/Pancarsuyu)
Ömer Faruk Çelik… See the full description on the dataset page: https://huggingface.co/datasets/barathanasln/turkish_llm_finetune_dataset_4_topics.collected-turkish-instructions-v0.1This dataset is the result of merging and cleaning data from the following sources:
Turkish Poems Cleaned
Turkish Reading Comprehension Question Answering Dataset
Stanford ALPaCA Cleaned Turkish Translated
Turkish Poems
Turkish Folk Song Lyrics
The data has been merged and processed for quality and consistency to create this dataset.
Turkish-QA
Aixr Türkçe Veri Seti
Aixr tarafından oluşturulan bu veri seti, Türkçe kaynak arayanlar için hazırlanmıştır. Yazılım geliştirme, radyoloji, tıbbi görüntüleme ve diğer konularda bilgi sağlayan bu veri seti, öğrenme süreçlerini kolaylaştırmayı ve Türkçe dilinde yapay zeka geliştirme süreçlerini desteklemeyi hedefler.
Veri Seti Hakkında
Amaç:
Bu veri seti, Türkçe içerikler arayan araştırmacılar, geliştiriciler ve eğitimciler için bir kaynak olarak tasarlanmıştır.
İçerik… See the full description on the dataset page: https://huggingface.co/datasets/Aixr/Turkish-QA.turkish-grammar-mmlu
Turkish-Grammar-MMLU
This dataset, created by Turkish-DB, is a multiple-choice question-answering (QA) dataset covering Turkish grammar topics. It is designed to evaluate model performance on various Turkish grammar subjects, similar to the MMLU (Massive Multitask Language Understanding) benchmark.
Overview
Name: Turkish-Grammar-MMLU
Provider: Turkish-DB
Task: Multiple-Choice QA
Modality: Text
Format: CSV (also accessible via API in Parquet format)
Language: Turkish… See the full description on the dataset page: https://huggingface.co/datasets/turkish-db/turkish-grammar-mmlu.turkish-university-mevzuat
Turkey University Regulation Data Collection
This dataset provides a comprehensive collection of regulatory documents of Turkish universities obtained from mevzuat.gov.tr.
It includes full texts of regulations with detailed publication information and unique identifiers.
Overview
Data Sources: mevzuat.gov.tr website
Technologies Used: Selenium, BeautifulSoup, Python
Data Formats: CSV
CSV Data Structure
Column
Description
Üniversite
Name of the… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/turkish-university-mevzuat.nvidia-nemotron-model-reasoning-dataset-turkish
Nemotron Reasoning Challenge - Turkish
Turkish translation of the training data from NVIDIA's Nemotron Model Reasoning Challenge
Each row is a reasoning puzzle framed in an "Alice's Wonderland" setting. Given a few input/output examples, the model needs to figure out the hidden rule and apply it to a new input.
Category
Rows
Description
bit
1602
Hidden bit manipulation rule on 8-bit binary numbers
grav
1597
Falling distance with a modified gravitational constant… See the full description on the dataset page: https://huggingface.co/datasets/mramazan/nvidia-nemotron-model-reasoning-dataset-turkish.Turkish-STEM-DPO-Dataset
Turkish STEM DPO Dataset
Dataset Summary
The Turkish STEM DPO (Direct Preference Optimization) dataset is a comprehensive synthetic resource containing 16,177 high-quality preference pairs designed to enhance the reasoning capabilities of Turkish language models in mathematics, physics, and programming.
The dataset leverages a preference-based learning approach: each instance pairs a carefully crafted, expert-level solution with a deliberately flawed or incomplete… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-STEM-DPO-Dataset.Turkish-STEM-DPO-Dataset
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti yusufbaykaloglu tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: yusufbaykaloglu/Turkish-STEM-DPO-Dataset
🔗 Derleyen Platform: VeriPazarı
Türkçe STEM DPO Veri Seti (Turkish STEM DPO Dataset)
Veri Seti Özeti
Turkish STEM DPO (Doğrudan Tercih… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/Turkish-STEM-DPO-Dataset.
