CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mrfg /turkish-court-decisions Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/mrfg/turkish-court-decisions.tabulartext-generation10M<n<100M5 likes1.3k downloads1mo agoHugging Face02hasankursun /turkish-corpus-100b Turkish Corpus 100B (TC-100B) Dataset Summary The Turkish Corpus 100B (TC-100B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Turkish. Comprising approximately 105 Billion tokens (measured with Qwen/Llama3 tokenizer), it represents one of the largest open resources for Turkish LLM pretraining. The dataset is engineered for a two-stage training pipeline: Pretrain Subset (~103B Tokens): A diverse mix of… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/turkish-corpus-100b.texttext-generation100M<n<1B8 likes854 downloads3mo agoHugging Face03AlicanKiraz0 /Turkish-SFT-Dataset-v1.0 Turkish-SFT-Dataset-v1.01 Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci 🔎 Özet Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.texttext-classification1K<n<10K52 likes467 downloads11mo agoHugging Face04boun-tabilab /turkish_parliamentary_data Grand National Assembly Corpus of Türkiye (GNACT) A comprehensive collection of Turkish parliamentary transcripts spanning over 100 years (1920–present), from 10 legislative bodies. Includes both Ottoman Turkish (1920–1928) and Modern Turkish (1928–present) texts. Loading the dataset from datasets import load_dataset # Strategy 1: full session documents, all bodies (default) ds = load_dataset("boun-tabilab/turkish_parliamentary_data", "full_sessions", split="train") #… See the full description on the dataset page: https://huggingface.co/datasets/boun-tabilab/turkish_parliamentary_data.tabulartext-generation1M<n<10M8 likes415 downloads6mo agoHugging Face05AYueksel /TurkishMMLU TurkishMMLU This repository contains Code and Data Analysis of TurkishMMLU for ACL 24 SIGTURK Workshop. TurkishMMLU is a multiple-choice dataset for Turkish Natural Language Processing (NLP) community based on Turkish Highschool Curricula for nine subjects. To access this dataset please send an email to: arda.yueksel@tum.de or akoksal@cis.lmu.de. Abstract Multiple choice question answering tasks evaluate the reasoning, comprehension, and mathematical abilities of… See the full description on the dataset page: https://huggingface.co/datasets/AYueksel/TurkishMMLU.textquestion-answering1K<n<10K17 likes356 downloads2y agoHugging Face06Alptekinege /turkish-court-decisions Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Alptekinege/turkish-court-decisions.tabulartext-generation10M<n<100M3 likes348 downloads24d agoHugging Face07Gyrevortex /turkish-court-decisions Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Gyrevortex/turkish-court-decisions.tabulartext-generation10M<n<100M2 likes259 downloads1mo agoHugging Face08serdarsrts /turkish-court-decisions-duplicate Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-court-decisions-duplicate.tabulartext-generation10M<n<100M1 likes231 downloads28d agoHugging Face09NoirZangetsu /Flutter-Code-with-Questions-Dataset-Turkish Flutter Code with Questions Dataset (Turkish) 📦 Dataset Name: flutter_code_with_questions Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir. 📁 Dataset Format Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.textquestion-answering1K<n<10K0 likes217 downloads2mo agoHugging Face10turkish-nlp-suite /InstrucTurca InstrucTurca v1.0.0 is a diverse synthetic instruction tuning dataset crafted for instruction-tuning Turkish LLMs. The data is compiled data various English datasets and sources, such as code instructions, poems, summarized texts, medical texts, and more. Dataset content BI55/MedText checkai/instruction-poems garage-bAInd/Open-Platypus Locutusque/ColumnedChatCombined nampdn-ai/tiny-codes Open-Orca/OpenOrca pubmed_qa TIGER-Lab/MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/InstrucTurca.texttext-generation1M<n<10M40 likes198 downloads2y agoHugging Face11GoktugD /turkish-extractive-qa-1.5m Turkish Extractive QA 1.5M v2 Cevap metni ve başlangıç konumu doğrulanabilir Türkçe çıkarımsal soru-cevap kayıtları. Doğrulanmış boyut Train: 1,470,000 Validation: 15,000 Test: 15,000 Toplam: 1,500,000 Ana görev sütunları: id, context, question, answer, answer_start, question_type Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-extractive-qa-1.5m.tabularquestion-answering1M<n<10M0 likes190 downloads2mo agoHugging Face12bugrabilge /Bilge-Turkish-CoT-50K Bilge: Turkish Chain-of-Thought Dataset (50K) 50,000 örneklik Türkçe Chain-of-Thought (CoT) reasoning fine-tuning veri seti. Bilge, Türkçe büyük dil modellerinin adım adım düşünme (reasoning) kapasitesini geliştirmek amacıyla hazırlanmış bir Chain-of-Thought veri setidir. Veri setindeki her örnek, modelin önce <think> blokları içinde görünür bir muhakeme süreci yürütmesini, ardından kullanıcıya yapılandırılmış ve detaylı bir cevap vermesini öğretmek üzere tasarlanmıştır. Bu… See the full description on the dataset page: https://huggingface.co/datasets/bugrabilge/Bilge-Turkish-CoT-50K.texttext-generation10K<n<100K9 likes182 downloads4mo agoHugging Face13CtnkyaABC /turkish-law-corpus ⚖️ Turkish Law — 106 Kanun Korpusu & Soru-Cevap106 Statutes Corpus & QA 🇹🇷 Türk hukukunun en çok kullanılan 106 kanunu, madde madde temizlenmiş 16.001 metin parçası ve bu maddelere dayalı 5.011 Türkçe soru-cevap çifti. Tamamı resmî kaynaktan (mevzuat.gov.tr), RAG ve yapay zekâ uygulamaları için hazır. 🇬🇧 The 106 most widely used Turkish statutes as 16,001 clean, article-level text chunks, plus 5,011 Turkish question-answer pairs grounded in those articles. All from the… See the full description on the dataset page: https://huggingface.co/datasets/CtnkyaABC/turkish-law-corpus.textquestion-answering10K<n<100K3 likes181 downloads2mo agoHugging Face14tunahanf /turkish-medicine-law turkish-medicine-law Bu veri seti, Türkçe tıp ve sağlık hukuku alanında hazırlanmıştır. Türkçe hukuk alanında genel amaçlı birkaç kaynak bulunuyor, ama tıp hukuku özelinde hazırlanmış bir veri seti şimdiye kadar yoktu. Bu proje o boşluğu doldurmayı amaçlıyor. Veri setindeki örnekler hukukçular, bilirkişiler ve sağlık kuruluşlarının hukuk birimleri için hazırlandı. Hastaya veya hekime doğrudan hukuki görüş sunmak amacıyla kullanılmak üzere tasarlanmadı. Buradaki çıktılar bir ön… See the full description on the dataset page: https://huggingface.co/datasets/tunahanf/turkish-medicine-law.texttext-generation1K<n<10K0 likes177 downloads10d agoHugging Face15umarigan /PD12M-TurkishTranslated from English to Tuskish language from: https://huggingface.co/datasets/Spawning/PD12M One of the biggest text-to-image dataset in Turkish language Metadata The metadata is made available through a series of parquet files with the following schema: text: Translated caption for the image. id: A unique identifier for the image. url: The URL of the image. caption: A caption for the image. width: The width of the image in pixels. height: The height of the image in pixels.… See the full description on the dataset page: https://huggingface.co/datasets/umarigan/PD12M-Turkish.imagequestion-answering10M<n<100M8 likes171 downloads2y agoHugging Face16sedayzc /turkish-medical-rag 🩺 Turkish Medical RAG Hierarchical Parent–Child Retrieval-Augmented Generation for Turkish Medical Documents 📌 Proje Hakkında Bu proje, Türkçe tıbbi dokümanlar üzerinde çalışan uçtan uca bir Retrieval-Augmented Generation (RAG) sistemi geliştirmek amacıyla hazırlanmıştır. Sistem bir kullanıcı sorusu aldığında önce doküman koleksiyonundaki küçük ve anlamsal olarak odaklı parçalar (child chunks)… See the full description on the dataset page: https://huggingface.co/datasets/sedayzc/turkish-medical-rag.tabularquestion-answering1K<n<10K0 likes161 downloads2mo agoHugging Face17emirms /turkish-competition-authority-decisions Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026 The complete published decision history of the Turkish Competition Authority (Rekabet Kurumu) — every Competition Board decision the regulator has made public, in full text, with derived structural metadata. 10,367 decisions · 113,297 pages · 323 million characters · 29 years Every decision carries its outcome, the articles of Law 4054 it turns on, the panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/emirms/turkish-competition-authority-decisions.tabulartext-classification10K<n<100K3 likes149 downloads1mo agoHugging Face18logicBombExe /turkish_cyber_security_controls_benchmark Turkish Cyber Security Controls Benchmark Türkçe siber güvenlik kontrol seçimi ve kontrol denetimi yeteneğini ölçmek için hazırlanmış, senaryo tabanlı çoktan seçmeli değerlendirme kümesidir. v0.1.0, uzman incelemesine açık ilk sürümdür ve NIST SP 800-53 Rev. 5, Release 5.2.0 kontrol kataloğunu hedefler. Kapsam 100 Türkçe senaryo NIST SP 800-53'ün 20 kontrol ailesinin her birinden 5 soru 64 kontrol seçimi sorusu 17 denetim kanıtı sorusu 19 denetim yargısı sorusu… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/turkish_cyber_security_controls_benchmark.textquestion-answeringn<1K4 likes146 downloads2mo agoHugging Face19AlicanKiraz0 /Turkish-Finance-SFT-Dataset 🇹🇷 Turkish Finance SFT Dataset Türkçe Finans Alanına Özel Supervised Fine-Tuning (SFT) Dataseti 📋 Dataset Özeti Bu dataset, Türkçe finans asistanı LLM'lerin eğitimi için özel olarak tasarlanmış, kapsamlı bir Supervised Fine-Tuning (SFT) veri setidir. Kripto para, borsa, teknik analiz, temel analiz, risk yönetimi ve finansal regülasyonlar dahil olmak üzere geniş bir yelpazede yaklaşık 10 milyon token boyutunda soru-cevap çifti verisi içermektedir. Dataset, hem… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-Finance-SFT-Dataset.textquestion-answering1K<n<10K63 likes144 downloads8mo agoHugging Face20sixfingerdev /turkish-qa-multi-dialog-dataset Turkish QA & Multi-Dialog Dataset Bu depo, iki farklı Türkçe veri kaynağının birleştirilmiş ve temizlenmiş sürümünü içerir: Yaklaşık 19.000 adet soru-cevap (QA) örneği Çok adımlı, doğal Türkçe sohbetlerden oluşan diyalog verileri Bu dataset, hem genel amaçlı Türkçe QA modelleri hem de sohbet/chatbot modelleri için uygundur. Veri İçeriği QA Bölümü (~19K) SQuAD benzeri yapıdan dönüştürülmüş input–output örnekleri Her satır: tek bir soru ve net bir cevap… See the full description on the dataset page: https://huggingface.co/datasets/sixfingerdev/turkish-qa-multi-dialog-dataset.textquestion-answering10K<n<100K4 likes140 downloads10mo agoHugging Face21AlicanKiraz0 /Turkish-CoT-Instruct-Dataset 🇹🇷 Turkish CoT Instruct Dataset Türkçe Düşünme Zinciri (Chain-of-Thought) İçeren Talimat Veri Seti Bu veri seti, modellerin Türkçe adım adım akıl yürütme (reasoning) yeteneğini geliştirmek için hazırlanmıştır. Her örnekte model, cevabı vermeden önce <think> ... </think> etiketleri arasında tamamen Türkçe olarak adım adım düşünür, ardından ayrıntılı bir nihai cevap sunar (DeepSeek-R1 tarzı biçim). Örnek sayısı: 4.868 Dil: Türkçe Biçim: Sohbet (messages) — system / user /… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-CoT-Instruct-Dataset.texttext-generation1K<n<10K20 likes136 downloads2mo agoHugging Face22ituperceptron /turkish_medical_reasoning Türkçe Medikal Reasoning Veri Seti Bu veri seti FreedomIntelligence/medical-o1-verifiable-problem veri setinin Türkçeye çevirilmiş bir alt kümesidir. Çevirdiğimiz veri seti 7,208 satır içermektedir. Veri setinde bulunan sütunlar aşağıda açıklanmıştır: question: Medikal soruların bulunduğu sütun. answer_content: DeepSeek-R1 modeli tarafından oluşturulmuş İngilizce yanıtların Türkçeye çevrilmiş hali.* reasoning_content: DeepSeek-R1 modeli tarafından oluşturulmuş İngilizce akıl… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/turkish_medical_reasoning.textquestion-answering1K<n<10K20 likes128 downloads8mo agoHugging Face23erenfazlioglu /turkish-instruct-reasoning-dpo-3.4m 🇹🇷 Turkish Instruct · Reasoning · DPO — ~3.4M The largest open native-Turkish instruction-tuning suite: SFT + chain-of-thought reasoning + DPO preferences, with an independently verified reasoning tier and a unique Turkey-grounded slice. En büyük açık native Türkçe talimat-eğitim seti: SFT + adım-adım muhakeme (CoT) + DPO tercih çiftleri; bağımsız doğrulanmış muhakeme katmanı ve Türkiye-temelli özgün dilim içerir. 📦 Examples ~3.44M ( SFT 3.26M · DPO 181k )… See the full description on the dataset page: https://huggingface.co/datasets/erenfazlioglu/turkish-instruct-reasoning-dpo-3.4m.texttext-generation1M<n<10M2 likes126 downloads3mo agoHugging Face24alibayram /Bilge-Turkish-CoT-50K Bilge: Turkish Chain-of-Thought Dataset (50K) 50,000 örneklik Türkçe Chain-of-Thought (CoT) reasoning fine-tuning veri seti. Bilge, Türkçe büyük dil modellerinin adım adım düşünme (reasoning) kapasitesini geliştirmek amacıyla hazırlanmış bir Chain-of-Thought veri setidir. Veri setindeki her örnek, modelin önce <think> blokları içinde görünür bir muhakeme süreci yürütmesini, ardından kullanıcıya yapılandırılmış ve detaylı bir cevap vermesini öğretmek üzere tasarlanmıştır. Bu… See the full description on the dataset page: https://huggingface.co/datasets/alibayram/Bilge-Turkish-CoT-50K.texttext-generation10K<n<100K1 likes124 downloads4mo agoHugging Face25bysismo /Turkish-instruction-3m_Soru_Cevap 🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance). 🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/Turkish-instruction-3m_Soru_Cevap.textquestion-answering1M<n<10M3 likes103 downloads1mo agoHugging Face26erythropygia /ThinkingData-200K-Turkish Dataset Card for ThinkingData-200K-Turkish Language: Turkish Dataset Description This repository contains a dataset for Turkish version of the Deepseek 1.5B model. The translation was performed using the Google translation model to ensure high-quality, accurate translation. Dataset Details Size: ≈205K Translation tool: Google Translate Data format: Prompt, Think, Response texttext-generation100K<n<1M16 likes98 downloads2y agoHugging Face27emre /TARA_Turkish_LLM_Benchmark TARA: Turkish Advanced Reasoning Assessment Veri Seti *Img Credit: Open AI ChatGPT **English version is given below.** Evaluation Notebook / Değerlendirme Not Defteri Dataset Summary TARA (Turkish Advanced Reasoning Assessment), Türkçe dilindeki Büyük Dil Modellerinin (LLM'ler) gelişmiş akıl yürütme yeteneklerini çoklu alanlarda ölçmek için tasarlanmış, zorluk derecesine göre sınıflandırılmış bir benchmark veri setidir. Bu veri seti, LLM'lerin sadece bilgi… See the full description on the dataset page: https://huggingface.co/datasets/emre/TARA_Turkish_LLM_Benchmark.textquestion-answeringn<1K28 likes90 downloads1y agoHugging Face28Taklaxbr /turkish_law_qa_dataset Not: Bu veri seti orijinal olarak OrionCAF tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir. 🔗 Orijinal Kaynak: OrionCAF/turkish_law_qa_dataset 🔗 Derleyen Platform: VeriPazarı 📚 Türkçe Hukuk Soru-Cevap Veri Seti (Turkish Law QA Dataset) Turkish Law QA Dataset, Türk hukuku üzerine odaklanmış, çeşitli hukuki metinlerden, içtihatlardan ve mevzuatlardan titizlikle derlenmiş 18.300+ soru-cevap çiftinden oluşan kapsamlı bir veri setidir.… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/turkish_law_qa_dataset.textquestion-answering10K<n<100K0 likes88 downloads4mo agoHugging Face29Renicames /turkish-law-chatbot MindLaw için Hukuk Veri Seti Bu veri seti, MindLaw modelinin eğitimi için oluşturulmuş olup, Türkçe hukuk alanına özgü metinlerden derlenmiştir. Veri seti, anayasanın sunduğu içeriklerden ve anayasayı açıklayan hukuki metinlerden oluşmaktadır. Ayrıca, bireylerin avukatlara sıkça yönlendirebilecekleri sorular formatında düzenlenmiş hukuki sorular ve cevapları da içermektedir. Veri Seti İçeriği Anayasa Metinleri: Türkiye Cumhuriyeti Anayasası'nın çeşitli maddeleri ve… See the full description on the dataset page: https://huggingface.co/datasets/Renicames/turkish-law-chatbot.textquestion-answering10K<n<100K31 likes81 downloads2y agoHugging Face30DevHunterAI /turkish-reasoning-data Turkish Reasoning Data A large-scale synthetic Turkish logic-puzzle dataset designed for pretraining and fine-tuning language models on structured reasoning tasks. Dataset Summary Property Value Language Turkish (tr) Format Parquet (HuggingFace Datasets compatible) Approximate tokens ~100 million Number of examples ~720,000 Split train only License CC BY 4.0 Load with datasets from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/DevHunterAI/turkish-reasoning-data.texttext-generation100K<n<1M0 likes78 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.