CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tascib /turkish-llm-dataset Turkish Pretraining Corpus Dataset Description This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models. This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/tascib/turkish-llm-dataset.text100M<n<1B15 likes11k downloads5mo agoHugging Face02turkish-nlp-suite /TrGLUE TrGLUE - The First Non-Translate Natural Language Understanding Benchmark for Turkish Dataset Card for TrGLUE TrGLUE is a natural language understanding benchmarking dataset including several single sentence and sentence pair classification tasks. The inspiration is clearly the original GLUE benchmark. Tasks Single Sentence Tasks TrCOLA The original Corpus of Linguistic Acceptability consists of sentences compiled from English literature textbooks.… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/TrGLUE.texttext-classification100K<n<1M6 likes567 downloads9mo agoHugging Face03turkish-nlp-suite /temiz-OSCAR Dataset Card for Temiz OSCAR Temiz OSCAR is a corpora collection consisting of cleaned versions of original OSCAR corpora. This collection is made up of four datasets: OSCAR-2019, OSCAR-2109, OSCAR-2201 and OSCAR-2301 This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. Dataset num instances size num of words OSCAR-2019 3.671.430 7.7G 976M OSCAR-2109 8.472.809 18G 2.22B OSCAR-2201… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-OSCAR.textfill-mask10M<n<100M5 likes479 downloads11mo agoHugging Face04AlicanKiraz0 /Turkish-SFT-Dataset-v1.0 Turkish-SFT-Dataset-v1.01 Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci 🔎 Özet Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.texttext-classification1K<n<10K52 likes437 downloads11mo agoHugging Face05TFLai /Turkish-AlpacaStanford alpaca turkish: Stanford Alpaca text10K<n<100K28 likes434 downloads3y agoHugging Face06AYueksel /TurkishMMLU TurkishMMLU This repository contains Code and Data Analysis of TurkishMMLU for ACL 24 SIGTURK Workshop. TurkishMMLU is a multiple-choice dataset for Turkish Natural Language Processing (NLP) community based on Turkish Highschool Curricula for nine subjects. To access this dataset please send an email to: arda.yueksel@tum.de or akoksal@cis.lmu.de. Abstract Multiple choice question answering tasks evaluate the reasoning, comprehension, and mathematical abilities of… See the full description on the dataset page: https://huggingface.co/datasets/AYueksel/TurkishMMLU.textquestion-answering1K<n<10K17 likes402 downloads2y agoHugging Face07turkish-nlp-suite /AkademikDerlem Dataset Card for AkademikDerlem AkademikDerlem is a scientific text corpus for Turkish, gathered from misc academical publication websites. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. This collection is made up of five datasets: Articles, Academic-Abstracts, Medical-Articles, Medical-Abstracts, and Bilkent-Writings. The Bilkent-Writings dataset comes from creative writings produced in the… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/AkademikDerlem.textfill-mask100K<n<1M6 likes250 downloads11mo agoHugging Face08AltaySec /turkish-llm-injection 🇹🇷 AltaySec Turkish LLM Prompt Injection Dataset (v0.2) Türkiye'nin ilk Türkçe-öncelikli, kategorize edilmiş LLM prompt injection veri seti — genişletilmiş sürüm. 📌 TL;DR 300 elle/üretim-destekli hazırlanmış Türkçe prompt injection payload'u, 12 saldırı kategorisi × 25, OWASP LLM Top 10 (2025) ile eşlenmiş. v0.1'in 120 çekirdek payload'una, AltayDuel arenasındaki bulgular ışığında üretilip düşmanca kalite/dedup denetiminden geçirilmiş 180 yeni payload eklendi.… See the full description on the dataset page: https://huggingface.co/datasets/AltaySec/turkish-llm-injection.texttext-classificationn<1K2 likes248 downloads1mo agoHugging Face09turkish-nlp-suite /Havadis Dataset Card for Havadis Havadis is a high quality and large Turkish news corpus, indeed the largest Turkish news corpus ever. This corpus is scraped from online news sebsites and includes text from popular newspapers such as CNN Türk Habertürk Hürriyet Millyet NTV Posta Sabah Star Sözcü Takvim . The instances are first crawled from the corresponding websites, then went throught an extensive cleaning process. We eliminated instances that are too short, too repetetive… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Havadis.textfill-mask100K<n<1M6 likes233 downloads3mo agoHugging Face10turkish-nlp-suite /TurkishHateMap Turkish Hate Map - A Large Scale and Diverse Hate Speech Dataset for Turkish Dataset Summary Turkish Hate Map (TuHaMa for short) is a big scale Turkish hate speech dataset that includes diverse target groups such as misogyny, political animosity, animal aversion, vegan antipathy, ethnic group hostility, and more. The dataset includes a total of 52K instances with 13 target groups. The dataset includes 4 labels, offensive, hate, neutral and civilized. Here is the… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/TurkishHateMap.texttext-classification10K<n<100K4 likes212 downloads2y agoHugging Face11VoidOaz /turkish-llm-dataset Turkish Pretraining Corpus Dataset Description This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models. This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/VoidOaz/turkish-llm-dataset.text100M<n<1B1 likes205 downloads13d agoHugging Face12turkish-nlp-suite /InstrucTurca InstrucTurca v1.0.0 is a diverse synthetic instruction tuning dataset crafted for instruction-tuning Turkish LLMs. The data is compiled data various English datasets and sources, such as code instructions, poems, summarized texts, medical texts, and more. Dataset content BI55/MedText checkai/instruction-poems garage-bAInd/Open-Platypus Locutusque/ColumnedChatCombined nampdn-ai/tiny-codes Open-Orca/OpenOrca pubmed_qa TIGER-Lab/MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/InstrucTurca.texttext-generation1M<n<10M40 likes204 downloads2y agoHugging Face13moganai /turkish-resmi-gazete Turkish Resmi Gazete (Official Gazette) — Full Corpus Türkiye Resmî Gazete'sinin 2019-01-02'den 2026-07-17'ye kadar yayımlanan tüm sayılarının tam metnini içeren bir derlemdir. Kanunlar, yönetmelikler, Cumhurbaşkanı kararları, tebliğler, Anayasa Mahkemesi kararları, kurul kararları ve atama kararnameleri dahil olmak üzere Resmî Gazete'de o tarih aralığında yayımlanmış her türden resmî belge yer almaktadır. İçerik ve Kaynak Resmî Gazete, günlük yayınlarını… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkish-resmi-gazete.texttext-generation10K<n<100K1 likes204 downloads2d agoHugging Face14turkish-nlp-suite /ForumSohbetleri Dataset Card for ForumSohbetleri ForumSohbetleri a web forum tetx corpus for Turkish, indeed first large-scale Turkish forum text corpus. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. This collection is made up of several subsets, each subset is gathered from the corresponding forum website. Forum websites contains diverse topics, ladies only, tech, economics, life, relations and much more...… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/ForumSohbetleri.textfill-mask1M<n<10M5 likes201 downloads11mo agoHugging Face15tascib /turkish-instruction Turkish Instruction Dataset Dataset Description This dataset is a large-scale Turkish instruction-tuning corpus created by combining multiple publicly available datasets and applying cleaning and deduplication steps. It is designed for training and evaluating large language models (LLMs) in Turkish. The dataset was prepared as part of a capstone project by students from Sabancı University. Data Sources The dataset is constructed from the following sources:… See the full description on the dataset page: https://huggingface.co/datasets/tascib/turkish-instruction.text100K<n<1M2 likes196 downloads6mo agoHugging Face16tunahanf /turkish-medicine-law turkish-medicine-law Bu veri seti, Türkçe tıp ve sağlık hukuku alanında hazırlanmıştır. Türkçe hukuk alanında genel amaçlı birkaç kaynak bulunuyor, ama tıp hukuku özelinde hazırlanmış bir veri seti şimdiye kadar yoktu. Bu proje o boşluğu doldurmayı amaçlıyor. Veri setindeki örnekler hukukçular, bilirkişiler ve sağlık kuruluşlarının hukuk birimleri için hazırlandı. Hastaya veya hekime doğrudan hukuki görüş sunmak amacıyla kullanılmak üzere tasarlanmadı. Buradaki çıktılar bir ön… See the full description on the dataset page: https://huggingface.co/datasets/tunahanf/turkish-medicine-law.texttext-generation1K<n<10K0 likes183 downloads10d agoHugging Face17bugrabilge /Bilge-Turkish-CoT-50K Bilge: Turkish Chain-of-Thought Dataset (50K) 50,000 örneklik Türkçe Chain-of-Thought (CoT) reasoning fine-tuning veri seti. Bilge, Türkçe büyük dil modellerinin adım adım düşünme (reasoning) kapasitesini geliştirmek amacıyla hazırlanmış bir Chain-of-Thought veri setidir. Veri setindeki her örnek, modelin önce <think> blokları içinde görünür bir muhakeme süreci yürütmesini, ardından kullanıcıya yapılandırılmış ve detaylı bir cevap vermesini öğretmek üzere tasarlanmıştır. Bu… See the full description on the dataset page: https://huggingface.co/datasets/bugrabilge/Bilge-Turkish-CoT-50K.texttext-generation10K<n<100K9 likes182 downloads4mo agoHugging Face18orhunc /Bias-Evaluation-TurkishTranslation of bias evaluation framework of May et al. (2019) from this repository and this paper into Turkish. There is a total of 37 tests including tests addressing gender-bias as well as tests designed to evaluate the ethnic bias toward Kurdish people in Türkiye context. Abstract of the paper: While the growing size of pre-trained language models has led to large improvements in a variety of natural language processing tasks, the success of these models comes with a price: They are trained… See the full description on the dataset page: https://huggingface.co/datasets/orhunc/Bias-Evaluation-Turkish.textn<1K1 likes177 downloads4y agoHugging Face19finansai /kap-turkish-financial-sentiment KAP Turkish Financial Sentiment Dataset Türkçe KAP (Kamuyu Aydınlatma Platformu) bildirimleri için çok boyutlu finansal analiz dataseti. Dataset Bilgileri Özellik Değer Kayıt Sayısı 3,839 Dil Türkçe Kaynak KAP Bildirimleri Etiketleme GPT-4 (Teacher Model) Format JSONL (Chat Messages) Kullanım Alanları Türkçe finansal sentiment analizi KAP bildirimi sınıflandırma Volatilite tahmini İlişkili taraf işlemi tespiti LLM fine-tuning (Qwen… See the full description on the dataset page: https://huggingface.co/datasets/finansai/kap-turkish-financial-sentiment.texttext-classification1K<n<10K1 likes166 downloads10mo agoHugging Face20ytu-ce-cosmos /Turkish-LLaVA-Finetune 🔥 TurkishLLaVA Finetuning Dataset This repository contains the dataset used for finetuning the Turkish-LLaVA-v0.1 model. The finetuning process was performed using this dataset, which was concatenated with Turkish-Books to enhance the model's performance. The details of this dataset, along with the finetuning results, will be shared in our upcoming paper (Soon..). Finetuning Configuration During the finetuning phase, both the projection matrix and the language model… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/Turkish-LLaVA-Finetune.text100K<n<1M6 likes160 downloads1y agoHugging Face21logicBombExe /turkish_cyber_security_controls_benchmark Turkish Cyber Security Controls Benchmark Türkçe siber güvenlik kontrol seçimi ve kontrol denetimi yeteneğini ölçmek için hazırlanmış, senaryo tabanlı çoktan seçmeli değerlendirme kümesidir. v0.1.0, uzman incelemesine açık ilk sürümdür ve NIST SP 800-53 Rev. 5, Release 5.2.0 kontrol kataloğunu hedefler. Kapsam 100 Türkçe senaryo NIST SP 800-53'ün 20 kontrol ailesinin her birinden 5 soru 64 kontrol seçimi sorusu 17 denetim kanıtı sorusu 19 denetim yargısı sorusu… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/turkish_cyber_security_controls_benchmark.textquestion-answeringn<1K4 likes150 downloads2mo agoHugging Face22turkish-nlp-datasets /eksisozluk-ekonomi-ve-finans-tr Ekşi Sözlük Türkçe Teknoloji Dataset Ekşi Sözlük'ten derlenen, teknoloji kategorisine ait Türkçe kullanıcı entry'lerinden oluşan bir veri setidir. Türkçe NLP araştırmaları ve LLM eğitimi için hazırlanmıştır. İçerik Makroekonomik kavramlar, yatırım araçları, kripto para ve güncel ekonomik gelişmeleri kapsayan 39 farklı başlık altında toplanmış entry'lerden oluşmaktadır. Alan Kapsanan Başlıklar Ekonomik Kriz & Genel Durum 2025-2026 ekonomik krizi, devalüasyon… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-datasets/eksisozluk-ekonomi-ve-finans-tr.text10K<n<100K0 likes144 downloads5mo agoHugging Face23logicBombExe /turkish_political_position_benchmark Turkish Political Position Benchmark The Turkish Political Position Benchmark measures how language models respond to normative statements about Turkish politics. It reports ideological dimension scores and response similarity to documented political-party reference profiles. The benchmark does not claim that a model belongs to a party, has a voting intention, or possesses political beliefs. A party similarity score only means that the model produced a similar pattern of answers… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/turkish_political_position_benchmark.tabulartext-classificationn<1K3 likes143 downloads2mo agoHugging Face24KocLab-Bilkent /turkish-constitutional-court Dataset Summary This dataset is extracted from the following Github repo, which was created for the journal paper with URL https://www.sciencedirect.com/science/article/abs/pii/S0306457321001692. https://github.com/koc-lab/law-turk The dataset includes 1290 court case decision texts from the Turkish Court of Cassation. Each sample has one label, which is the ruling of the court. The possible rulings are "Violation" and "No violation". There are 1290 samples. 1141 of these samples… See the full description on the dataset page: https://huggingface.co/datasets/KocLab-Bilkent/turkish-constitutional-court.texttext-classification1K<n<10K9 likes138 downloads4y agoHugging Face25AlicanKiraz0 /Turkish-Finance-SFT-Dataset 🇹🇷 Turkish Finance SFT Dataset Türkçe Finans Alanına Özel Supervised Fine-Tuning (SFT) Dataseti 📋 Dataset Özeti Bu dataset, Türkçe finans asistanı LLM'lerin eğitimi için özel olarak tasarlanmış, kapsamlı bir Supervised Fine-Tuning (SFT) veri setidir. Kripto para, borsa, teknik analiz, temel analiz, risk yönetimi ve finansal regülasyonlar dahil olmak üzere geniş bir yelpazede yaklaşık 10 milyon token boyutunda soru-cevap çifti verisi içermektedir. Dataset, hem… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-Finance-SFT-Dataset.textquestion-answering1K<n<10K63 likes138 downloads8mo agoHugging Face26sixfingerdev /turkish-qa-multi-dialog-dataset Turkish QA & Multi-Dialog Dataset Bu depo, iki farklı Türkçe veri kaynağının birleştirilmiş ve temizlenmiş sürümünü içerir: Yaklaşık 19.000 adet soru-cevap (QA) örneği Çok adımlı, doğal Türkçe sohbetlerden oluşan diyalog verileri Bu dataset, hem genel amaçlı Türkçe QA modelleri hem de sohbet/chatbot modelleri için uygundur. Veri İçeriği QA Bölümü (~19K) SQuAD benzeri yapıdan dönüştürülmüş input–output örnekleri Her satır: tek bir soru ve net bir cevap… See the full description on the dataset page: https://huggingface.co/datasets/sixfingerdev/turkish-qa-multi-dialog-dataset.textquestion-answering10K<n<100K4 likes135 downloads10mo agoHugging Face27AlicanKiraz0 /Turkish-CoT-Instruct-Dataset 🇹🇷 Turkish CoT Instruct Dataset Türkçe Düşünme Zinciri (Chain-of-Thought) İçeren Talimat Veri Seti Bu veri seti, modellerin Türkçe adım adım akıl yürütme (reasoning) yeteneğini geliştirmek için hazırlanmıştır. Her örnekte model, cevabı vermeden önce <think> ... </think> etiketleri arasında tamamen Türkçe olarak adım adım düşünür, ardından ayrıntılı bir nihai cevap sunar (DeepSeek-R1 tarzı biçim). Örnek sayısı: 4.868 Dil: Türkçe Biçim: Sohbet (messages) — system / user /… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-CoT-Instruct-Dataset.texttext-generation1K<n<10K20 likes133 downloads2mo agoHugging Face28beratcmn /turkish-prompt-injections Turkish Prompt Injections Translated version of deepset/prompt-injections. I highly recommend training a model with both translated and the original texts instead of just using only the translated prompts. I will also add more Turkish injection examples soon. texttext-classificationn<1K5 likes124 downloads3y agoHugging Face29ituperceptron /turkish_medical_reasoning Türkçe Medikal Reasoning Veri Seti Bu veri seti FreedomIntelligence/medical-o1-verifiable-problem veri setinin Türkçeye çevirilmiş bir alt kümesidir. Çevirdiğimiz veri seti 7,208 satır içermektedir. Veri setinde bulunan sütunlar aşağıda açıklanmıştır: question: Medikal soruların bulunduğu sütun. answer_content: DeepSeek-R1 modeli tarafından oluşturulmuş İngilizce yanıtların Türkçeye çevrilmiş hali.* reasoning_content: DeepSeek-R1 modeli tarafından oluşturulmuş İngilizce akıl… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/turkish_medical_reasoning.textquestion-answering1K<n<10K20 likes116 downloads8mo agoHugging Face30bysismo /Turkish-instruction-3m_Soru_Cevap 🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance). 🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/Turkish-instruction-3m_Soru_Cevap.textquestion-answering1M<n<10M3 likes111 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.