CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01moganai /turkishfineweb2-cleaned TurkishFineweb2-Cleaned A Turkish web corpus derived from the Turkish (tur_Latn) subset of FineWeb-2, augmented with an additional quality-classification layer and a near-duplicate removal pass. 📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum Source FineWeb-2 is a large-scale, multilingual web corpus built from Common Crawl. This dataset covers the Turkish (tur_Latn) portion of FineWeb-2, spanning the… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkishfineweb2-cleaned.tabulartext-generation10M<n<100M5 likes1.7k downloads2d agoHugging Face02Ethosoft /Turkish_corpus Turkish Corpus 🇹🇷 Turkish Corpus is a large-scale cleaned Turkish text dataset created by collecting public Turkish corpora and extracting Turkish-language portions from multilingual datasets. The dataset is designed for Turkish Natural Language Processing research, language model pretraining, tokenizer training, embedding models, retrieval systems, and general Turkish language understanding tasks. The main purpose of this dataset is to provide a practical, scalable, and… See the full description on the dataset page: https://huggingface.co/datasets/Ethosoft/Turkish_corpus.tabulartext-generation1M<n<10M1 likes1.4k downloads5mo agoHugging Face03mrfg /turkish-court-decisions Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/mrfg/turkish-court-decisions.tabulartext-generation10M<n<100M5 likes1.3k downloads1mo agoHugging Face04boun-tabilab /turkish_parliamentary_data Grand National Assembly Corpus of Türkiye (GNACT) A comprehensive collection of Turkish parliamentary transcripts spanning over 100 years (1920–present), from 10 legislative bodies. Includes both Ottoman Turkish (1920–1928) and Modern Turkish (1928–present) texts. Loading the dataset from datasets import load_dataset # Strategy 1: full session documents, all bodies (default) ds = load_dataset("boun-tabilab/turkish_parliamentary_data", "full_sessions", split="train") #… See the full description on the dataset page: https://huggingface.co/datasets/boun-tabilab/turkish_parliamentary_data.tabulartext-generation1M<n<10M8 likes397 downloads6mo agoHugging Face05Alptekinege /turkish-court-decisions Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Alptekinege/turkish-court-decisions.tabulartext-generation10M<n<100M3 likes326 downloads24d agoHugging Face06Gyrevortex /turkish-court-decisions Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Gyrevortex/turkish-court-decisions.tabulartext-generation10M<n<100M2 likes251 downloads1mo agoHugging Face07serdarsrts /turkish-court-decisions-duplicate Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-court-decisions-duplicate.tabulartext-generation10M<n<100M1 likes203 downloads28d agoHugging Face08Ba2han /fineweb-2-turkish-categorized-long altaidevorg/fineweb-2-turkish-categorized long filtered Turkish texts Source: altaidevorg/fineweb-2-turkish-categorized (config: default). The script streamed 10,000,000 raw source rows before stopping. Categories ads, adult content, sports, tabloid were rejected before length and quality filtering. Retained rows contain 3,000–16,500 characters and passed the iteration-5 Turkish language, repetition, glue-word, punctuation, SEO, and soft information-density filters. Selected… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/fineweb-2-turkish-categorized-long.tabulartext-generation100K<n<1M0 likes132 downloads2mo agoHugging Face09Ethosoft /Turkish-Recipe-Corpus Turkish Recipe Corpus (TRC-30K) Türkçe'nin en kapsamlı açık kaynaklı tarif veri seti.74,768 Türkçe tariften oluşan, 3 farklı NLP görevine hazır yapılandırılmış corpus. Ethosoft Research · huggingface.co/Ethosoft · ethosoft.org Dataset Özeti Değer Toplam tarif 74,768 Dil Türkçe (tr) Lisans CC BY 4.0 Konfigürasyonlar 3 (structured, instruction, ingredient2recipe) Split Train / Validation / Test (80 / 10 / 10) Ortalama adım sayısı 5.23 /… See the full description on the dataset page: https://huggingface.co/datasets/Ethosoft/Turkish-Recipe-Corpus.tabulartext-generation100K<n<1M1 likes107 downloads4mo agoHugging Face10furkankarli /turkish-brand-bias-evaluations Turkish Brand Bias Evaluations / Türkçe Marka Yanlılığı Değerlendirmeleri Furkan Karlı tarafından Türkçe ürün ve hizmet önerilerindeki marka görünürlüğünü incelemek amacıyla oluşturulmuş LLM değerlendirme veri setidir. An LLM evaluation dataset curated by Furkan Karlı to study brand visibility in Turkish product and service recommendations. Veri seti özeti 300 tamamlanmış ve judge edilmiş yanıt Domainler: VPN (150) ve kozmetik (150) Koşullar: web araması kapalı… See the full description on the dataset page: https://huggingface.co/datasets/furkankarli/turkish-brand-bias-evaluations.tabulartext-generationn<1K1 likes102 downloads21d agoHugging Face11bilalabic /turkish-tool-calling Türkçe Tool-Calling Veri Seti 56.247 kayıt. xLAM/APIGen 60k ve NVIDIA When2Call'dan türetilmiş, üç davranış sınıfı içeren Türkçe function-calling veri seti. from datasets import load_dataset ds = load_dataset("bilalabic/turkish-tool-calling") # mesaj listesi ds = load_dataset("bilalabic/turkish-tool-calling", "table") # düz tablo ds = load_dataset("bilalabic/turkish-tool-calling", "sharegpt") # ShareGPT İçerik Kayıt 56.247… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/turkish-tool-calling.tabulartext-generation100K<n<1M0 likes89 downloads2mo agoHugging Face12fatihburakkaragoz /old-nogay-turkish-ocr-corpus Old Nogay Turkish OCR Corpus This is a small OCR-derived corpus of historical Nogay Turkish / Turkic textual material, collected, classified, extracted, and packaged by Fatih Burak Karagöz / CDLI.ai for exploratory NLP, historical corpus work, and OCR-quality analysis. We are proud to release this as a contribution to Turkish NLP, Turkic-language NLP, historical NLP, and Turkish studies. The goal is not to pretend that a small OCR corpus is a polished benchmark. The goal is more… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/old-nogay-turkish-ocr-corpus.tabulartext-generationn<1K0 likes79 downloads5mo agoHugging Face13Quardo /Turkish-Chat_GPT-4O Quardo/Turkish-Chat_GPT-4O Description This is a simple dataset generated by OpenAI's GPT-4O (gpt-4o-2024-08-06). The dataset includes various entries created and evaluated by the AI model, providing a unique collection of Turkish chat data for analysis and research. Warning Please note that this dataset may contain errors or inconsistencies as it is fully generated by an AI model. It is highly recommended to check and edit the data before usage, as AI can… See the full description on the dataset page: https://huggingface.co/datasets/Quardo/Turkish-Chat_GPT-4O.tabulartext-generation1K<n<10K12 likes63 downloads2y agoHugging Face14berkbirkan /turkish-seo-reasoning-benchmark-results Turkish SEO Reasoning Benchmark Results Bu dataset, Turkish SEO Reasoning benchmark'ının altı farklı model/checkpoint üzerinde çalıştırılmış ham tahminlerini, metriklerini ve tekrar üretim manifestlerini içerir. Fine-tuned model: berkbirkan/gemma-3-1b-turkish-seo-reasoning-lora Sonuç Fine-tuned Gemma 3 1B modeli 22,23 skorla ilk sırada yer aldı. Aynı base model 11,96 skor elde etti. Mutlak artış: +10,28 puan Göreli artış: %85,97 Fine-tuned model hata sayısı:… See the full description on the dataset page: https://huggingface.co/datasets/berkbirkan/turkish-seo-reasoning-benchmark-results.tabulartext-generationn<1K0 likes54 downloads2mo agoHugging Face15erythropygia /Instruction-280K-Turkish Dataset Card for Instruction-280K-Turkish Language: Turkish Dataset Description This repository contains a dataset for Turkish version of the Deepseek 1.5B model. The translation was performed using the Google translation model to ensure high-quality, accurate translation. Dataset Details Size: ≈280K Translation tool: Google Translate Data format: Prompt, Response tabulartext-generation100K<n<1M1 likes49 downloads2y agoHugging Face16okg /turkish-poemsTurkish poems scraped from antoloji.com. Features consists of id, poet name, poem rating and the poem. tabulartext-generation1K<n<10K6 likes43 downloads4y agoHugging Face17mustafakemal0146 /Turkish-Lyric-Intelligence-v2 Turkish Lyric Intelligence v2 Turkish Lyric Intelligence v2 is a derived, research-oriented dataset for building controllable Turkish lyric generation and lyric-analysis systems. It adds section structure, orthographic syllable counts, rhyme-ending candidates, repetition signals, and review queues to the source Genius Turkish Dataset. This release is intended as an intermediate data layer. Automated rhyme, prosody, and task labels are candidate annotations, not expert literary… See the full description on the dataset page: https://huggingface.co/datasets/mustafakemal0146/Turkish-Lyric-Intelligence-v2.tabulartext-generation100K<n<1M0 likes43 downloads2mo agoHugging Face18Ba2han /mogan-turkish-web-long moganai/mogan-turkish-web long filtered Turkish texts Source: moganai/mogan-turkish-web (config: default, revision: d773a0efd1b7daf72c7909c83dbd385e7e3564d7). Rows contain 4,000–16,000 characters and passed the iteration-5 Turkish language, repetition, glue-word, punctuation, SEO, and soft information-density filters. Selected rows: 2,805,010. Generated by process_hf_dataset.py. See summary.json for counts and thresholds. tabulartext-generation1M<n<10M0 likes43 downloads5d agoHugging Face19beratcmn /rephrased-instruction-turkish-poems This is a rephrased version of my previous dataset beratcmn/instruction-turkish-poems. I used the same instructions but I rephrased them to be more clear and understandable also added more variety to the format. tabulartext-generation1K<n<10K6 likes36 downloads3y agoHugging Face20Taklaxbr /turkish-math-rlvr Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti barandinho tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir. 🔗 Orijinal Kaynak: barandinho/turkish-math-rlvr 🔗 Derleyen Platform: VeriPazarı Türkçe Matematiksel Akıl Yürütme (RLVR Eğitim Veri Seti) Bu Veri Seti Nedir? Bu veri seti, zayıf bir modelin başarı oranına… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/turkish-math-rlvr.tabulartext-generation1K<n<10K0 likes22 downloads3mo agoHugging Face21yusufbaykaloglu /Turkish-Legislation-DPO Turkish-Legislation-DPO Dataset Turkish-Legislation-DPO is a large-scale Direct Preference Optimization (DPO) training dataset specifically designed to align Turkish language models with expertise in the fields of law and regulation. Dataset Summary The Turkish-Legislation-DPO dataset contains 23,596 carefully selected preference pairs, all generated by google/gemma-2-2b-it model and improved through systematic quality assessment protocols. Each example consists of… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-Legislation-DPO.tabularquestion-answering10K<n<100K2 likes20 downloads1y agoHugging Face22GoktugD /turkish-number-words-1m Turkish Number Words 1M v2 0 ile 999.999 arasındaki her tamsayının Türkçe yazıyla karşılığı. Doğrulanmış boyut Train: 980,000 Validation: 10,000 Test: 10,000 Toplam: 1,000,000 Ana görev sütunları: id, number, words Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance, generator_version, generator_sha256… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-number-words-1m.tabulartext-generation1M<n<10M0 likes19 downloads2mo agoHugging Face23Lightcap /nocisnn-turkish-logic-traces NociSNN Turkish Mathematical Logic Traces This dataset records structured Turkish mathematical reasoning traces generated by ogulcanaydogan/Turkish-LLM-7B-Instruct-GGUF:Q4_K_M in a continuously running NociSNN evaluation loop. Data Generation Problems are generated from deterministic Turkish templates covering arithmetic, percentage, proportion and single-variable linear equations. Ground-truth answers are computed by the task generator rather than the language… See the full description on the dataset page: https://huggingface.co/datasets/Lightcap/nocisnn-turkish-logic-traces.tabulartext-generation1K<n<10K1 likes16 downloads4mo agoHugging Face24yusufbaykaloglu /Turkish-STEM-DPO-Dataset Turkish STEM DPO Dataset Dataset Summary The Turkish STEM DPO (Direct Preference Optimization) dataset is a comprehensive synthetic resource containing 16,177 high-quality preference pairs designed to enhance the reasoning capabilities of Turkish language models in mathematics, physics, and programming. The dataset leverages a preference-based learning approach: each instance pairs a carefully crafted, expert-level solution with a deliberately flawed or incomplete… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-STEM-DPO-Dataset.tabulartext-generation10K<n<100K4 likes12 downloads1y agoHugging Face25Taklaxbr /Turkish-STEM-DPO-Dataset Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti yusufbaykaloglu tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir. 🔗 Orijinal Kaynak: yusufbaykaloglu/Turkish-STEM-DPO-Dataset 🔗 Derleyen Platform: VeriPazarı Türkçe STEM DPO Veri Seti (Turkish STEM DPO Dataset) Veri Seti Özeti Turkish STEM DPO (Doğrudan Tercih… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/Turkish-STEM-DPO-Dataset.tabulartext-generation10K<n<100K0 likes11 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.