CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01moganai /turkishfineweb2-cleaned TurkishFineweb2-Cleaned A Turkish web corpus derived from the Turkish (tur_Latn) subset of FineWeb-2, augmented with an additional quality-classification layer and a near-duplicate removal pass. 📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum Source FineWeb-2 is a large-scale, multilingual web corpus built from Common Crawl. This dataset covers the Turkish (tur_Latn) portion of FineWeb-2, spanning the… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkishfineweb2-cleaned.tabulartext-generation10M<n<100M5 likes1.7k downloads2d agoHugging Face02Ethosoft /Turkish_corpus Turkish Corpus 🇹🇷 Turkish Corpus is a large-scale cleaned Turkish text dataset created by collecting public Turkish corpora and extracting Turkish-language portions from multilingual datasets. The dataset is designed for Turkish Natural Language Processing research, language model pretraining, tokenizer training, embedding models, retrieval systems, and general Turkish language understanding tasks. The main purpose of this dataset is to provide a practical, scalable, and… See the full description on the dataset page: https://huggingface.co/datasets/Ethosoft/Turkish_corpus.tabulartext-generation1M<n<10M1 likes1.4k downloads5mo agoHugging Face03mrfg /turkish-court-decisions Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/mrfg/turkish-court-decisions.tabulartext-generation10M<n<100M5 likes1.3k downloads1mo agoHugging Face04moganai /mogan-turkish-web Mogan Turkish Web A large-scale Turkish web corpus derived from monthly Common Crawl snapshots covering the period from January 2025 to June 2026. The corpus was constructed by extracting Turkish-language content from raw Common Crawl WARC/WET dumps, followed by language filtering, PII masking, and near-duplicate removal. 📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum Dataset Summary This dataset… See the full description on the dataset page: https://huggingface.co/datasets/moganai/mogan-turkish-web.texttext-generation10M<n<100M6 likes1.2k downloads2d agoHugging Face05hasankursun /turkish-corpus-100b Turkish Corpus 100B (TC-100B) Dataset Summary The Turkish Corpus 100B (TC-100B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Turkish. Comprising approximately 105 Billion tokens (measured with Qwen/Llama3 tokenizer), it represents one of the largest open resources for Turkish LLM pretraining. The dataset is engineered for a two-stage training pipeline: Pretrain Subset (~103B Tokens): A diverse mix of… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/turkish-corpus-100b.texttext-generation100M<n<1B8 likes846 downloads3mo agoHugging Face063nesdeniz /turkish-daily-dialogues-5k Turkish Daily Dialogues 5K Exactly 5,000 synthetic, multi-turn Turkish conversations covering ordinary daily-life situations. The corpus is designed as a small, auditable baseline for dialogue modelling, instruction-format experiments, augmentation research, and Turkish-language evaluation—not as a substitute for conversations written by real people. Provenance in one sentence: the Turkish source scenario library was drafted with AI assistance specifically for this project… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/turkish-daily-dialogues-5k.texttext-generation1K<n<10K2 likes510 downloads2mo agoHugging Face07turkish-nlp-suite /temiz-OSCAR Dataset Card for Temiz OSCAR Temiz OSCAR is a corpora collection consisting of cleaned versions of original OSCAR corpora. This collection is made up of four datasets: OSCAR-2019, OSCAR-2109, OSCAR-2201 and OSCAR-2301 This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. Dataset num instances size num of words OSCAR-2019 3.671.430 7.7G 976M OSCAR-2109 8.472.809 18G 2.22B OSCAR-2201… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-OSCAR.textfill-mask10M<n<100M5 likes479 downloads11mo agoHugging Face08AlicanKiraz0 /Turkish-SFT-Dataset-v1.0 Turkish-SFT-Dataset-v1.01 Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci 🔎 Özet Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.texttext-classification1K<n<10K52 likes437 downloads11mo agoHugging Face09boun-tabilab /turkish_parliamentary_data Grand National Assembly Corpus of Türkiye (GNACT) A comprehensive collection of Turkish parliamentary transcripts spanning over 100 years (1920–present), from 10 legislative bodies. Includes both Ottoman Turkish (1920–1928) and Modern Turkish (1928–present) texts. Loading the dataset from datasets import load_dataset # Strategy 1: full session documents, all bodies (default) ds = load_dataset("boun-tabilab/turkish_parliamentary_data", "full_sessions", split="train") #… See the full description on the dataset page: https://huggingface.co/datasets/boun-tabilab/turkish_parliamentary_data.tabulartext-generation1M<n<10M8 likes397 downloads6mo agoHugging Face10Alptekinege /turkish-court-decisions Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Alptekinege/turkish-court-decisions.tabulartext-generation10M<n<100M3 likes326 downloads24d agoHugging Face11Gyrevortex /turkish-court-decisions Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Gyrevortex/turkish-court-decisions.tabulartext-generation10M<n<100M2 likes251 downloads1mo agoHugging Face12turkish-nlp-suite /AkademikDerlem Dataset Card for AkademikDerlem AkademikDerlem is a scientific text corpus for Turkish, gathered from misc academical publication websites. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. This collection is made up of five datasets: Articles, Academic-Abstracts, Medical-Articles, Medical-Abstracts, and Bilkent-Writings. The Bilkent-Writings dataset comes from creative writings produced in the… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/AkademikDerlem.textfill-mask100K<n<1M6 likes250 downloads11mo agoHugging Face13AltaySec /turkish-llm-injection 🇹🇷 AltaySec Turkish LLM Prompt Injection Dataset (v0.2) Türkiye'nin ilk Türkçe-öncelikli, kategorize edilmiş LLM prompt injection veri seti — genişletilmiş sürüm. 📌 TL;DR 300 elle/üretim-destekli hazırlanmış Türkçe prompt injection payload'u, 12 saldırı kategorisi × 25, OWASP LLM Top 10 (2025) ile eşlenmiş. v0.1'in 120 çekirdek payload'una, AltayDuel arenasındaki bulgular ışığında üretilip düşmanca kalite/dedup denetiminden geçirilmiş 180 yeni payload eklendi.… See the full description on the dataset page: https://huggingface.co/datasets/AltaySec/turkish-llm-injection.texttext-classificationn<1K2 likes248 downloads1mo agoHugging Face14turkish-nlp-suite /Havadis Dataset Card for Havadis Havadis is a high quality and large Turkish news corpus, indeed the largest Turkish news corpus ever. This corpus is scraped from online news sebsites and includes text from popular newspapers such as CNN Türk Habertürk Hürriyet Millyet NTV Posta Sabah Star Sözcü Takvim . The instances are first crawled from the corresponding websites, then went throught an extensive cleaning process. We eliminated instances that are too short, too repetetive… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Havadis.textfill-mask100K<n<1M6 likes233 downloads3mo agoHugging Face15NoirZangetsu /Flutter-Code-with-Questions-Dataset-Turkish Flutter Code with Questions Dataset (Turkish) 📦 Dataset Name: flutter_code_with_questions Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir. 📁 Dataset Format Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.textquestion-answering1K<n<10K0 likes228 downloads2mo agoHugging Face16turkish-nlp-suite /InstrucTurca InstrucTurca v1.0.0 is a diverse synthetic instruction tuning dataset crafted for instruction-tuning Turkish LLMs. The data is compiled data various English datasets and sources, such as code instructions, poems, summarized texts, medical texts, and more. Dataset content BI55/MedText checkai/instruction-poems garage-bAInd/Open-Platypus Locutusque/ColumnedChatCombined nampdn-ai/tiny-codes Open-Orca/OpenOrca pubmed_qa TIGER-Lab/MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/InstrucTurca.texttext-generation1M<n<10M40 likes204 downloads2y agoHugging Face17moganai /turkish-resmi-gazete Turkish Resmi Gazete (Official Gazette) — Full Corpus Türkiye Resmî Gazete'sinin 2019-01-02'den 2026-07-17'ye kadar yayımlanan tüm sayılarının tam metnini içeren bir derlemdir. Kanunlar, yönetmelikler, Cumhurbaşkanı kararları, tebliğler, Anayasa Mahkemesi kararları, kurul kararları ve atama kararnameleri dahil olmak üzere Resmî Gazete'de o tarih aralığında yayımlanmış her türden resmî belge yer almaktadır. İçerik ve Kaynak Resmî Gazete, günlük yayınlarını… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkish-resmi-gazete.texttext-generation10K<n<100K1 likes204 downloads2d agoHugging Face18serdarsrts /turkish-court-decisions-duplicate Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-court-decisions-duplicate.tabulartext-generation10M<n<100M1 likes203 downloads28d agoHugging Face19turkish-nlp-suite /ForumSohbetleri Dataset Card for ForumSohbetleri ForumSohbetleri a web forum tetx corpus for Turkish, indeed first large-scale Turkish forum text corpus. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. This collection is made up of several subsets, each subset is gathered from the corresponding forum website. Forum websites contains diverse topics, ladies only, tech, economics, life, relations and much more...… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/ForumSohbetleri.textfill-mask1M<n<10M5 likes201 downloads11mo agoHugging Face20tunahanf /turkish-medicine-law turkish-medicine-law Bu veri seti, Türkçe tıp ve sağlık hukuku alanında hazırlanmıştır. Türkçe hukuk alanında genel amaçlı birkaç kaynak bulunuyor, ama tıp hukuku özelinde hazırlanmış bir veri seti şimdiye kadar yoktu. Bu proje o boşluğu doldurmayı amaçlıyor. Veri setindeki örnekler hukukçular, bilirkişiler ve sağlık kuruluşlarının hukuk birimleri için hazırlandı. Hastaya veya hekime doğrudan hukuki görüş sunmak amacıyla kullanılmak üzere tasarlanmadı. Buradaki çıktılar bir ön… See the full description on the dataset page: https://huggingface.co/datasets/tunahanf/turkish-medicine-law.texttext-generation1K<n<10K0 likes183 downloads10d agoHugging Face21bugrabilge /Bilge-Turkish-CoT-50K Bilge: Turkish Chain-of-Thought Dataset (50K) 50,000 örneklik Türkçe Chain-of-Thought (CoT) reasoning fine-tuning veri seti. Bilge, Türkçe büyük dil modellerinin adım adım düşünme (reasoning) kapasitesini geliştirmek amacıyla hazırlanmış bir Chain-of-Thought veri setidir. Veri setindeki her örnek, modelin önce <think> blokları içinde görünür bir muhakeme süreci yürütmesini, ardından kullanıcıya yapılandırılmış ve detaylı bir cevap vermesini öğretmek üzere tasarlanmıştır. Bu… See the full description on the dataset page: https://huggingface.co/datasets/bugrabilge/Bilge-Turkish-CoT-50K.texttext-generation10K<n<100K9 likes182 downloads4mo agoHugging Face22finansai /kap-turkish-financial-sentiment KAP Turkish Financial Sentiment Dataset Türkçe KAP (Kamuyu Aydınlatma Platformu) bildirimleri için çok boyutlu finansal analiz dataseti. Dataset Bilgileri Özellik Değer Kayıt Sayısı 3,839 Dil Türkçe Kaynak KAP Bildirimleri Etiketleme GPT-4 (Teacher Model) Format JSONL (Chat Messages) Kullanım Alanları Türkçe finansal sentiment analizi KAP bildirimi sınıflandırma Volatilite tahmini İlişkili taraf işlemi tespiti LLM fine-tuning (Qwen… See the full description on the dataset page: https://huggingface.co/datasets/finansai/kap-turkish-financial-sentiment.texttext-classification1K<n<10K1 likes166 downloads10mo agoHugging Face23Tuguberk /turkish-hermes-function-calling turkish-hermes-function-calling NousResearch/hermes-function-calling-v1 datasetinin Türkçe çevirisi — Hermes 2 Pro modelinin araç kullanımı ve yapılandırılmış çıktı yeteneklerini kazandıran orijinal veri seti. Genel Bakış Satır sayısı 11.567 Dil Türkçe (tr) Lisans Apache 2.0 Kaynak dataset NousResearch/hermes-function-calling-v1 Çeviri modeli DeepSeek V4 Flash (deepseek-chat) Ort. tur / konuşma 5,6 Çok turlu konuşma 6.120 (%52,9) Araç… See the full description on the dataset page: https://huggingface.co/datasets/Tuguberk/turkish-hermes-function-calling.texttext-generation10K<n<100K6 likes154 downloads4mo agoHugging Face24alibayram /turkish_mmlugated Turkish MMLU: Yapay Zeka ve Akademik Uygulamalar İçin En Kapsamlı ve Özgün Türkçe Veri Seti Önemli Not: Bu veri setini kullananların, özellikle Zenodo üzerinden alıntı yapmaları büyük önem taşımaktadır. Zenodo üzerinden yapılan alıntılar, veri setimizin bilimsel olarak daha geniş bir çevrede tanınmasını ve indekslenmesini sağlayacaktır. Lütfen aşağıdaki Zenodo DOI numarasını kullanarak veri setine atıfta bulunun: @dataset{bayram_2024_13378019, author = {Bayram, M. Ali}… See the full description on the dataset page: https://huggingface.co/datasets/alibayram/turkish_mmlu.texttext-generation100K<n<1M70 likes148 downloads1y agoHugging Face25GoktugD /turkish-punctuation-restoration-500k Turkish Punctuation Restoration 500K v2 Noktalama ve büyük harfleri kaldırılmış girişler ile hedef cümle çiftleri. Doğrulanmış boyut Train: 490,000 Validation: 5,000 Test: 5,000 Toplam: 500,000 Ana görev sütunları: id, unpunctuated_text, punctuated_text Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-punctuation-restoration-500k.texttext-generation100K<n<1M0 likes143 downloads2mo agoHugging Face26erenfazlioglu /turkish-instruct-reasoning-dpo-3.4m 🇹🇷 Turkish Instruct · Reasoning · DPO — ~3.4M The largest open native-Turkish instruction-tuning suite: SFT + chain-of-thought reasoning + DPO preferences, with an independently verified reasoning tier and a unique Turkey-grounded slice. En büyük açık native Türkçe talimat-eğitim seti: SFT + adım-adım muhakeme (CoT) + DPO tercih çiftleri; bağımsız doğrulanmış muhakeme katmanı ve Türkiye-temelli özgün dilim içerir. 📦 Examples ~3.44M ( SFT 3.26M · DPO 181k )… See the full description on the dataset page: https://huggingface.co/datasets/erenfazlioglu/turkish-instruct-reasoning-dpo-3.4m.texttext-generation1M<n<10M2 likes138 downloads3mo agoHugging Face27sixfingerdev /turkish-qa-multi-dialog-dataset Turkish QA & Multi-Dialog Dataset Bu depo, iki farklı Türkçe veri kaynağının birleştirilmiş ve temizlenmiş sürümünü içerir: Yaklaşık 19.000 adet soru-cevap (QA) örneği Çok adımlı, doğal Türkçe sohbetlerden oluşan diyalog verileri Bu dataset, hem genel amaçlı Türkçe QA modelleri hem de sohbet/chatbot modelleri için uygundur. Veri İçeriği QA Bölümü (~19K) SQuAD benzeri yapıdan dönüştürülmüş input–output örnekleri Her satır: tek bir soru ve net bir cevap… See the full description on the dataset page: https://huggingface.co/datasets/sixfingerdev/turkish-qa-multi-dialog-dataset.textquestion-answering10K<n<100K4 likes135 downloads10mo agoHugging Face28AlicanKiraz0 /Turkish-CoT-Instruct-Dataset 🇹🇷 Turkish CoT Instruct Dataset Türkçe Düşünme Zinciri (Chain-of-Thought) İçeren Talimat Veri Seti Bu veri seti, modellerin Türkçe adım adım akıl yürütme (reasoning) yeteneğini geliştirmek için hazırlanmıştır. Her örnekte model, cevabı vermeden önce <think> ... </think> etiketleri arasında tamamen Türkçe olarak adım adım düşünür, ardından ayrıntılı bir nihai cevap sunar (DeepSeek-R1 tarzı biçim). Örnek sayısı: 4.868 Dil: Türkçe Biçim: Sohbet (messages) — system / user /… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-CoT-Instruct-Dataset.texttext-generation1K<n<10K20 likes133 downloads2mo agoHugging Face29Ba2han /fineweb-2-turkish-categorized-long altaidevorg/fineweb-2-turkish-categorized long filtered Turkish texts Source: altaidevorg/fineweb-2-turkish-categorized (config: default). The script streamed 10,000,000 raw source rows before stopping. Categories ads, adult content, sports, tabloid were rejected before length and quality filtering. Retained rows contain 3,000–16,500 characters and passed the iteration-5 Turkish language, repetition, glue-word, punctuation, SEO, and soft information-density filters. Selected… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/fineweb-2-turkish-categorized-long.tabulartext-generation100K<n<1M0 likes132 downloads2mo agoHugging Face30selimfirat /bilkent-turkish-writings-dataset Compilation of Bilkent Turkish Writings Dataset Dataset Description This is a comprehensive compilation of Turkish creative writings from Bilkent University's Turkish 101 and Turkish 102 courses (2014-2025). The dataset contains 9119 student writings originally created by students and instructors, focusing on creativity, content, composition, grammar, spelling, and punctuation development. Note: This dataset is a compilation and digitization of publicly available writings… See the full description on the dataset page: https://huggingface.co/datasets/selimfirat/bilkent-turkish-writings-dataset.texttext-generation10K<n<100K19 likes127 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.