CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01turkish-nlp-suite /temiz-OSCAR Dataset Card for Temiz OSCAR Temiz OSCAR is a corpora collection consisting of cleaned versions of original OSCAR corpora. This collection is made up of four datasets: OSCAR-2019, OSCAR-2109, OSCAR-2201 and OSCAR-2301 This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. Dataset num instances size num of words OSCAR-2019 3.671.430 7.7G 976M OSCAR-2109 8.472.809 18G 2.22B OSCAR-2201… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-OSCAR.textfill-mask10M<n<100M5 likes479 downloads11mo agoHugging Face02AlicanKiraz0 /Turkish-SFT-Dataset-v1.0 Turkish-SFT-Dataset-v1.01 Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci 🔎 Özet Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.texttext-classification1K<n<10K52 likes437 downloads11mo agoHugging Face03turkish-nlp-suite /AkademikDerlem Dataset Card for AkademikDerlem AkademikDerlem is a scientific text corpus for Turkish, gathered from misc academical publication websites. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. This collection is made up of five datasets: Articles, Academic-Abstracts, Medical-Articles, Medical-Abstracts, and Bilkent-Writings. The Bilkent-Writings dataset comes from creative writings produced in the… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/AkademikDerlem.textfill-mask100K<n<1M6 likes250 downloads11mo agoHugging Face04AltaySec /turkish-llm-injection 🇹🇷 AltaySec Turkish LLM Prompt Injection Dataset (v0.2) Türkiye'nin ilk Türkçe-öncelikli, kategorize edilmiş LLM prompt injection veri seti — genişletilmiş sürüm. 📌 TL;DR 300 elle/üretim-destekli hazırlanmış Türkçe prompt injection payload'u, 12 saldırı kategorisi × 25, OWASP LLM Top 10 (2025) ile eşlenmiş. v0.1'in 120 çekirdek payload'una, AltayDuel arenasındaki bulgular ışığında üretilip düşmanca kalite/dedup denetiminden geçirilmiş 180 yeni payload eklendi.… See the full description on the dataset page: https://huggingface.co/datasets/AltaySec/turkish-llm-injection.texttext-classificationn<1K2 likes248 downloads1mo agoHugging Face05turkish-nlp-suite /Havadis Dataset Card for Havadis Havadis is a high quality and large Turkish news corpus, indeed the largest Turkish news corpus ever. This corpus is scraped from online news sebsites and includes text from popular newspapers such as CNN Türk Habertürk Hürriyet Millyet NTV Posta Sabah Star Sözcü Takvim . The instances are first crawled from the corresponding websites, then went throught an extensive cleaning process. We eliminated instances that are too short, too repetetive… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Havadis.textfill-mask100K<n<1M6 likes233 downloads2mo agoHugging Face06turkish-nlp-suite /InstrucTurca InstrucTurca v1.0.0 is a diverse synthetic instruction tuning dataset crafted for instruction-tuning Turkish LLMs. The data is compiled data various English datasets and sources, such as code instructions, poems, summarized texts, medical texts, and more. Dataset content BI55/MedText checkai/instruction-poems garage-bAInd/Open-Platypus Locutusque/ColumnedChatCombined nampdn-ai/tiny-codes Open-Orca/OpenOrca pubmed_qa TIGER-Lab/MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/InstrucTurca.texttext-generation1M<n<10M40 likes204 downloads2y agoHugging Face07moganai /turkish-resmi-gazete Turkish Resmi Gazete (Official Gazette) — Full Corpus Türkiye Resmî Gazete'sinin 2019-01-02'den 2026-07-17'ye kadar yayımlanan tüm sayılarının tam metnini içeren bir derlemdir. Kanunlar, yönetmelikler, Cumhurbaşkanı kararları, tebliğler, Anayasa Mahkemesi kararları, kurul kararları ve atama kararnameleri dahil olmak üzere Resmî Gazete'de o tarih aralığında yayımlanmış her türden resmî belge yer almaktadır. İçerik ve Kaynak Resmî Gazete, günlük yayınlarını… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkish-resmi-gazete.texttext-generation10K<n<100K1 likes204 downloads2d agoHugging Face08turkish-nlp-suite /ForumSohbetleri Dataset Card for ForumSohbetleri ForumSohbetleri a web forum tetx corpus for Turkish, indeed first large-scale Turkish forum text corpus. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. This collection is made up of several subsets, each subset is gathered from the corresponding forum website. Forum websites contains diverse topics, ladies only, tech, economics, life, relations and much more...… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/ForumSohbetleri.textfill-mask1M<n<10M5 likes201 downloads11mo agoHugging Face09tunahanf /turkish-medicine-law turkish-medicine-law Bu veri seti, Türkçe tıp ve sağlık hukuku alanında hazırlanmıştır. Türkçe hukuk alanında genel amaçlı birkaç kaynak bulunuyor, ama tıp hukuku özelinde hazırlanmış bir veri seti şimdiye kadar yoktu. Bu proje o boşluğu doldurmayı amaçlıyor. Veri setindeki örnekler hukukçular, bilirkişiler ve sağlık kuruluşlarının hukuk birimleri için hazırlandı. Hastaya veya hekime doğrudan hukuki görüş sunmak amacıyla kullanılmak üzere tasarlanmadı. Buradaki çıktılar bir ön… See the full description on the dataset page: https://huggingface.co/datasets/tunahanf/turkish-medicine-law.texttext-generation1K<n<10K0 likes183 downloads9d agoHugging Face10bugrabilge /Bilge-Turkish-CoT-50K Bilge: Turkish Chain-of-Thought Dataset (50K) 50,000 örneklik Türkçe Chain-of-Thought (CoT) reasoning fine-tuning veri seti. Bilge, Türkçe büyük dil modellerinin adım adım düşünme (reasoning) kapasitesini geliştirmek amacıyla hazırlanmış bir Chain-of-Thought veri setidir. Veri setindeki her örnek, modelin önce <think> blokları içinde görünür bir muhakeme süreci yürütmesini, ardından kullanıcıya yapılandırılmış ve detaylı bir cevap vermesini öğretmek üzere tasarlanmıştır. Bu… See the full description on the dataset page: https://huggingface.co/datasets/bugrabilge/Bilge-Turkish-CoT-50K.texttext-generation10K<n<100K9 likes182 downloads4mo agoHugging Face11finansai /kap-turkish-financial-sentiment KAP Turkish Financial Sentiment Dataset Türkçe KAP (Kamuyu Aydınlatma Platformu) bildirimleri için çok boyutlu finansal analiz dataseti. Dataset Bilgileri Özellik Değer Kayıt Sayısı 3,839 Dil Türkçe Kaynak KAP Bildirimleri Etiketleme GPT-4 (Teacher Model) Format JSONL (Chat Messages) Kullanım Alanları Türkçe finansal sentiment analizi KAP bildirimi sınıflandırma Volatilite tahmini İlişkili taraf işlemi tespiti LLM fine-tuning (Qwen… See the full description on the dataset page: https://huggingface.co/datasets/finansai/kap-turkish-financial-sentiment.texttext-classification1K<n<10K1 likes166 downloads10mo agoHugging Face12sixfingerdev /turkish-qa-multi-dialog-dataset Turkish QA & Multi-Dialog Dataset Bu depo, iki farklı Türkçe veri kaynağının birleştirilmiş ve temizlenmiş sürümünü içerir: Yaklaşık 19.000 adet soru-cevap (QA) örneği Çok adımlı, doğal Türkçe sohbetlerden oluşan diyalog verileri Bu dataset, hem genel amaçlı Türkçe QA modelleri hem de sohbet/chatbot modelleri için uygundur. Veri İçeriği QA Bölümü (~19K) SQuAD benzeri yapıdan dönüştürülmüş input–output örnekleri Her satır: tek bir soru ve net bir cevap… See the full description on the dataset page: https://huggingface.co/datasets/sixfingerdev/turkish-qa-multi-dialog-dataset.textquestion-answering10K<n<100K4 likes135 downloads10mo agoHugging Face13AlicanKiraz0 /Turkish-CoT-Instruct-Dataset 🇹🇷 Turkish CoT Instruct Dataset Türkçe Düşünme Zinciri (Chain-of-Thought) İçeren Talimat Veri Seti Bu veri seti, modellerin Türkçe adım adım akıl yürütme (reasoning) yeteneğini geliştirmek için hazırlanmıştır. Her örnekte model, cevabı vermeden önce <think> ... </think> etiketleri arasında tamamen Türkçe olarak adım adım düşünür, ardından ayrıntılı bir nihai cevap sunar (DeepSeek-R1 tarzı biçim). Örnek sayısı: 4.868 Dil: Türkçe Biçim: Sohbet (messages) — system / user /… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-CoT-Instruct-Dataset.texttext-generation1K<n<10K20 likes133 downloads2mo agoHugging Face14ituperceptron /turkish_medical_reasoning Türkçe Medikal Reasoning Veri Seti Bu veri seti FreedomIntelligence/medical-o1-verifiable-problem veri setinin Türkçeye çevirilmiş bir alt kümesidir. Çevirdiğimiz veri seti 7,208 satır içermektedir. Veri setinde bulunan sütunlar aşağıda açıklanmıştır: question: Medikal soruların bulunduğu sütun. answer_content: DeepSeek-R1 modeli tarafından oluşturulmuş İngilizce yanıtların Türkçeye çevrilmiş hali.* reasoning_content: DeepSeek-R1 modeli tarafından oluşturulmuş İngilizce akıl… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/turkish_medical_reasoning.textquestion-answering1K<n<10K20 likes116 downloads8mo agoHugging Face15bysismo /Turkish-instruction-3m_Soru_Cevap 🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance). 🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/Turkish-instruction-3m_Soru_Cevap.textquestion-answering1M<n<10M3 likes111 downloads1mo agoHugging Face16barandinho /turkish-reasoning-distilled-sft Turkish Reasoning Distilled SFT This dataset contains Turkish reasoning SFT data for barandinho/qwen3.5-27b-tudum-dapo-50. The teacher model also received a small RL run, but this dataset is its main supervised fine-tuning data. It combines generated Turkish reasoning traces from DAPO math, WebInstruct, AceCode, and verifier-compatible IFEval-style instruction-following sources with verified teacher-SFT traces from DAPO math, OpenThoughts science, OpenThoughts code, and system-chat… See the full description on the dataset page: https://huggingface.co/datasets/barandinho/turkish-reasoning-distilled-sft.texttext-generation1M<n<10M0 likes111 downloads4mo agoHugging Face17ytu-ce-cosmos /turkish-flow-drafter-prompts GitHub repo · Technical blog · Model collection Turkish prompts for Chained-Flow drafter training Chat-templated Turkish prompts used to train and evaluate the Turkish Flow-Drafter checkpoints for Qwen/Qwen3.5-4B / 9B / 27B. Prompts only — no completions. A drafter is trained on the target model's own hidden states, so continuations are generated locally by running the target over these prompts. Nothing here is a model output. split rows what it is v1/… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/turkish-flow-drafter-prompts.texttext-generation10K<n<100K0 likes101 downloads10d agoHugging Face18ogulcanaydogan /Turkish-LLM-v10-Training Turkish LLM Training Dataset v10 A curated corpus of 144,022 Turkish instruction-completion pairs used to train the Turkish LLM Family. Dataset Description This dataset was created to address the scarcity of high-quality Turkish instruction-following data for language model fine-tuning. It covers a broad range of topics including: Science & Technology (physics, chemistry, biology, computer science) History & Geography (Turkish and world history, geography) General… See the full description on the dataset page: https://huggingface.co/datasets/ogulcanaydogan/Turkish-LLM-v10-Training.texttext-generation100K<n<1M3 likes95 downloads7mo agoHugging Face19selimaktas /turkish-flow-drafter-prompts Turkish prompts for Chained-Flow drafter training Chat-templated Turkish prompts used to train and evaluate the Turkish Flow-Drafter checkpoints for Qwen/Qwen3.5-4B / 9B / 27B. Prompts only — no completions. A drafter is trained on the target model's own hidden states, so continuations are generated locally by running the target over these prompts. Nothing here is a model output. split rows what it is v1/ 29,100 train + 300 holdout the mixture the released Turkish… See the full description on the dataset page: https://huggingface.co/datasets/selimaktas/turkish-flow-drafter-prompts.texttext-generation10K<n<100K0 likes95 downloads22d agoHugging Face20bilalabic /turkish-tool-calling Türkçe Tool-Calling Veri Seti 56.247 kayıt. xLAM/APIGen 60k ve NVIDIA When2Call'dan türetilmiş, üç davranış sınıfı içeren Türkçe function-calling veri seti. from datasets import load_dataset ds = load_dataset("bilalabic/turkish-tool-calling") # mesaj listesi ds = load_dataset("bilalabic/turkish-tool-calling", "table") # düz tablo ds = load_dataset("bilalabic/turkish-tool-calling", "sharegpt") # ShareGPT İçerik Kayıt 56.247… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/turkish-tool-calling.tabulartext-generation100K<n<1M0 likes89 downloads2mo agoHugging Face21furkanyllmz /kap-turkish-financial-sentiment KAP Turkish Financial Sentiment Dataset Türkçe KAP (Kamuyu Aydınlatma Platformu) bildirimleri için çok boyutlu finansal analiz dataseti. Dataset Bilgileri Özellik Değer Kayıt Sayısı 3,839 Dil Türkçe Kaynak KAP Bildirimleri Etiketleme GPT-4 (Teacher Model) Format JSONL (Chat Messages) Kullanım Alanları Türkçe finansal sentiment analizi KAP bildirimi sınıflandırma Volatilite tahmini İlişkili taraf işlemi tespiti LLM fine-tuning (Qwen… See the full description on the dataset page: https://huggingface.co/datasets/furkanyllmz/kap-turkish-financial-sentiment.texttext-classification1K<n<10K0 likes88 downloads10mo agoHugging Face22berhaan /Turkish-CodeAlpaca-20k 🇹🇷 Turkish CodeAlpaca-20k Turkish CodeAlpaca-20k, popüler CodeAlpaca-20k veri kümesinin Türkçe çevirisidir.Bu veri kümesi, Türkçe kodlama görevlerinde instruction-tuning yapmak isteyen modeller için hazırlanmıştır.Tüm “instruction–input–output” çiftleri, orijinal İngilizce versiyondan anlam koruması gözetilerek çevrilmiştir. 📚 Veri Kümesi Hakkında Toplam örnek sayısı: ~20.000 Format: JSON / Parquet Alanlar: instruction: Modelin ne yapması gerektiğini… See the full description on the dataset page: https://huggingface.co/datasets/berhaan/Turkish-CodeAlpaca-20k.texttext-generation10K<n<100K2 likes73 downloads11mo agoHugging Face23cturan /turkish-synthetic-corpus Turkish Synthetic Corpus A synthetic Turkish text corpus with 1,871,131 documents, designed for Turkish language model training. About Inspired by HuggingFaceTB/smollm-corpus. Questions and prompts were sourced from the SmolLM Corpus pipeline; a language model then generated localized Turkish responses and documents around them. All credit for the original corpus design and methodology goes to the HuggingFace SmolLM team. The resulting dataset covers a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/cturan/turkish-synthetic-corpus.texttext-generation1M<n<10M0 likes69 downloads6mo agoHugging Face24besmabakirci01 /turkish-poetry-instruction-dataset Turkish Poetry Instruction Dataset Türkçe şiir üretimi için Alpaca instruction formatında hazırlanmış fine-tuning dataseti. Unsloth + Qwen LoRA eğitimi (3. ödev) için tasarlanmıştır. Format Her satır: { "instruction": "Kullanıcının şiir isteği", "input": "", "output": "Üretilmesi beklenen şiir metni" } Dosyalar Dosya Açıklama train.jsonl Eğitim (%90) validation.jsonl Doğrulama (%5) test.jsonl Test (%5) Toplam ~4.961 örnek;… See the full description on the dataset page: https://huggingface.co/datasets/besmabakirci01/turkish-poetry-instruction-dataset.texttext-generation1K<n<10K0 likes69 downloads2mo agoHugging Face25turkish-nlp-suite /temiz-WikiA cleaned version of Turkish Wikipedia dataset. The soource is wikimedia/wikipedia repo. The text is cleaned throughoutly, first of all we eliminated text that are shorter than a predetermined threshold of words and characters. Then we normalized with NFKC, cleaned some non-ASCII chars and normalized whitespaces. textfill-mask100K<n<1M4 likes64 downloads7mo agoHugging Face26Quardo /Turkish-Chat_GPT-4O Quardo/Turkish-Chat_GPT-4O Description This is a simple dataset generated by OpenAI's GPT-4O (gpt-4o-2024-08-06). The dataset includes various entries created and evaluated by the AI model, providing a unique collection of Turkish chat data for analysis and research. Warning Please note that this dataset may contain errors or inconsistencies as it is fully generated by an AI model. It is highly recommended to check and edit the data before usage, as AI can… See the full description on the dataset page: https://huggingface.co/datasets/Quardo/Turkish-Chat_GPT-4O.tabulartext-generation1K<n<10K12 likes63 downloads2y agoHugging Face27mmkocak /turkish-recipes-175K Turkish Recipes 175K Türkçe yemek tarifi dataseti — 174.975 deduplike tarif kaydı. Halka açık Türkçe yemek tarifi kaynaklarından toplanmış; başlık, malzeme listesi, talimatlar, kategori, etiket, porsiyon, pişirme süreleri ve besin değerleri içerir. Türkçe büyük dil modellerinin (LLM) pretraining ve instruction-tuning'i için hazırlanmıştır. İçerik Split Satır Boyut train 153.978 ~417 MB validation 10.498 ~28 MB test 10.499 ~28 MB Toplam 174.975… See the full description on the dataset page: https://huggingface.co/datasets/mmkocak/turkish-recipes-175K.texttext-generation100K<n<1M0 likes59 downloads4mo agoHugging Face28cturan /lamba-turkish-sft Lamba Turkish SFT Dataset This is a Turkish Supervised Fine-Tuning (SFT) dataset. The topic distribution is largely aligned with the Turkish High School (Lise) curriculum. It contains a wide variety of examples designed to improve model capabilities in: Instruction following Summarization (Özetleme) Information extraction (Bilgi çıkarma) General problem solving I hope this dataset will be beneficial to the open-source and AI community. Disclaimer Since the vast… See the full description on the dataset page: https://huggingface.co/datasets/cturan/lamba-turkish-sft.textquestion-answering100K<n<1M0 likes58 downloads6mo agoHugging Face29hasankursun /turkish-legislation-corpus Turkish Legislation Dataset This dataset contains Turkish legislative texts (laws) scraped from the official Mevzuat Bilgi Sistemi (mevzuat.gov.tr). The raw data has been cleaned, stripped of unnecessary whitespace, and structured into the JSONL format. Dataset Contents The dataset includes the full text of laws currently in force in the Republic of Turkey, along with relevant metadata. It is suitable for NLP tasks in the legal domain, such as LLM fine-tuning, RAG… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/turkish-legislation-corpus.texttext-classificationn<1K2 likes57 downloads10mo agoHugging Face30Uunan /turkish-planning-sft Turkish Planning SFT Turkish Planning SFT is a large-scale synthetic instruction-following dataset designed to improve the planning capabilities of Turkish Large Language Models (LLMs). Rather than focusing on factual question answering, the dataset teaches models how to transform user goals, requirements, and constraints into structured, practical, and actionable plans. The dataset is intended for Supervised Fine-Tuning (SFT) and follows a conversation-oriented format… See the full description on the dataset page: https://huggingface.co/datasets/Uunan/turkish-planning-sft.texttext-generation100K<n<1M0 likes55 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.