CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tascib /turkish-llm-dataset Turkish Pretraining Corpus Dataset Description This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models. This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/tascib/turkish-llm-dataset.text100M<n<1B15 likes11k downloads5mo agoHugging Face02serda-dev /turkish-raw-text-cleaned Turkish Raw Text Cleaned turkish-raw-text-cleaned, Türkçe dil modeli çalışmaları için hazırlanmış temizlenmiş ham metin veri kümesidir. Veri kümesi, turkish-nlp-suite çatısı altında yayımlanan Türkçe metin kaynaklarının temizlenmesi, filtrelenmesi ve model eğitimine daha uygun hale getirilmesiyle oluşturulmuştur. Bu çalışma özellikle Türkçe LLM ön-eğitimi, continual pre-training (CPT), tokenizer analizi, embedding modeli eğitimi, alan bağımsız Türkçe metin modelleme ve veri… See the full description on the dataset page: https://huggingface.co/datasets/serda-dev/turkish-raw-text-cleaned.text-generation1M<n<10M0 likes10k downloads3mo agoHugging Face03zinderud /risale-sohbet-turkish-2audio1K<n<10K0 likes3.6k downloads1y agoHugging Face04altaidevorg /fineweb-2-turkish-categorized What is this THis is the categorized version of the Turkish subset of the fineweb-2 dataset. It is an ongoing effort, and the details will be added soon with the rest of the dataset. tabular10M<n<100M15 likes3.5k downloads2y agoHugging Face05turkish-nlp-suite /BellaTurca Dataset Card for BellaTurca BellaTurca is the first large-scale Turkish corpus collection for training Turkish language models. The total size is around 245GB and 30 billion words. BellaTurca's focus is high quality, diversity as well as the size. This collection is made up of five datasets: AkademikDerlem, OzenliDerlem, ForumSohbetleri, Temiz OSCAR and Temiz mC4. Originally there was a book corpus included, but it is excluded due to containing copyrighted material. AkademikDerlem… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/BellaTurca.text10M<n<100M17 likes3.5k downloads7mo agoHugging Face06moganai /turkishfineweb2-cleaned TurkishFineweb2-Cleaned A Turkish web corpus derived from the Turkish (tur_Latn) subset of FineWeb-2, augmented with an additional quality-classification layer and a near-duplicate removal pass. 📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum Source FineWeb-2 is a large-scale, multilingual web corpus built from Common Crawl. This dataset covers the Turkish (tur_Latn) portion of FineWeb-2, spanning the… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkishfineweb2-cleaned.tabulartext-generation10M<n<100M5 likes1.7k downloads2d agoHugging Face07bysismo /Turkish-Python-instruction 🚀 DİKKAT VERİ SETİ GÜNCELLENME SÜRECİNE ALINMIŞTIR LÜTFEN AÇIKLAMAYI OKUYUNUZ. Turkish Python & System Engineering Dataset (BYSISMO v2.0) 25 Kategorilik Büyük Türkçe Python & Sistem Mühendisliği Havuzu 📢 SÜRÜM & DOĞRULAMA DURUMU (VERSION ROADMAP) v1.0 (Eski Arşiv - 289K / 8 Kategori): Yüksek kalite standartlarımız gereği yeniden yapılandırmaya alınmış ve dondurulmuştur. v2.0 (Yeni Master Sürüm - 416K+ / 17 Kategori): Kodlar yalnızca sözdizimi… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/Turkish-Python-instruction.text-generation100K<n<1M3 likes1.6k downloads8d agoHugging Face08Ethosoft /Turkish_corpus Turkish Corpus 🇹🇷 Turkish Corpus is a large-scale cleaned Turkish text dataset created by collecting public Turkish corpora and extracting Turkish-language portions from multilingual datasets. The dataset is designed for Turkish Natural Language Processing research, language model pretraining, tokenizer training, embedding models, retrieval systems, and general Turkish language understanding tasks. The main purpose of this dataset is to provide a practical, scalable, and… See the full description on the dataset page: https://huggingface.co/datasets/Ethosoft/Turkish_corpus.tabulartext-generation1M<n<10M1 likes1.4k downloads5mo agoHugging Face09mrfg /turkish-court-decisions Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/mrfg/turkish-court-decisions.tabulartext-generation10M<n<100M5 likes1.3k downloads1mo agoHugging Face10selmanbaysan /turkish_embedding_model_training_datatextsentence-similarity100M<n<1B5 likes1.3k downloads1y agoHugging Face11moganai /mogan-turkish-web Mogan Turkish Web A large-scale Turkish web corpus derived from monthly Common Crawl snapshots covering the period from January 2025 to June 2026. The corpus was constructed by extracting Turkish-language content from raw Common Crawl WARC/WET dumps, followed by language filtering, PII masking, and near-duplicate removal. 📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum Dataset Summary This dataset… See the full description on the dataset page: https://huggingface.co/datasets/moganai/mogan-turkish-web.texttext-generation10M<n<100M6 likes1.2k downloads2d agoHugging Face12hasankursun /turkish-corpus-100b Turkish Corpus 100B (TC-100B) Dataset Summary The Turkish Corpus 100B (TC-100B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Turkish. Comprising approximately 105 Billion tokens (measured with Qwen/Llama3 tokenizer), it represents one of the largest open resources for Turkish LLM pretraining. The dataset is engineered for a two-stage training pipeline: Pretrain Subset (~103B Tokens): A diverse mix of… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/turkish-corpus-100b.texttext-generation100M<n<1B8 likes846 downloads3mo agoHugging Face13Taklaxbr /turkish-tts-combined-raw Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından düzenlenmiştir. Orijinal veri seti afkfatih tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir. 🔗 Orijinal Kaynak: afkfatih/turkish-tts-combined-raw 🔗 Derleyen Platform: VeriPazarı Türkçe TTS Birleşik Veri Seti (Turkish TTS Combined) 7 farklı açık kaynak Türkçe TTS (Metinden Sese) veri setinin birleşimidir. ~81.500 örnek |… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/turkish-tts-combined-raw.audiotext-to-speech10K<n<100K0 likes791 downloads4mo agoHugging Face14merve /turkish_instructionstext10K<n<100K64 likes715 downloads3y agoHugging Face15turkish-nlp-suite /OzenliDerlem Dataset Card for OzenliDerlem OzenliDerlem (a.k.a CraftedCrawl) is a carefully assembled collection of top-notch web crawl data from handpicked websites, featuring articles, journals, and magazines. It focuses on gathering rich and detailed text content, especially longer articles. The collection covers a wide range of topics, including travel, news, culture, fairy tales and folklore, movie reviews, popular science, product and service complaints, fashion and self-care, trendy… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/OzenliDerlem.text1M<n<10M11 likes712 downloads7mo agoHugging Face16winvoker /turkish-sentiment-analysis-dataset Dataset This dataset contains positive , negative and notr sentences from several data sources given in the references. In the most sentiment models , there are only two labels; positive and negative. However , user input can be totally notr sentence. For such cases there were no data I could find. Therefore I created this dataset with 3 class. Positive and negative sentences are listed below. Notr examples are extraced from turkish wiki dump. In addition, added some random text… See the full description on the dataset page: https://huggingface.co/datasets/winvoker/turkish-sentiment-analysis-dataset.texttext-classification100K<n<1M49 likes677 downloads3y agoHugging Face17turkish-nlp-suite /temiz-mC4 Dataset Card for Temiz mC4 Temiz mC4 is the cleaned version of CulturaX corpus' Turkish split. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. split num instances size num of words train 76.432.893 168GB 21.06B Total 76.432.893 168GB 21.06B This collection includes web text, crawled from internet for mC4 corpus. CulturaX is even refined version of mC4 with quality filtering and… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-mC4.fill-mask2 likes572 downloads11mo agoHugging Face18turkish-nlp-suite /TrGLUE TrGLUE - The First Non-Translate Natural Language Understanding Benchmark for Turkish Dataset Card for TrGLUE TrGLUE is a natural language understanding benchmarking dataset including several single sentence and sentence pair classification tasks. The inspiration is clearly the original GLUE benchmark. Tasks Single Sentence Tasks TrCOLA The original Corpus of Linguistic Acceptability consists of sentences compiled from English literature textbooks.… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/TrGLUE.texttext-classification100K<n<1M6 likes567 downloads9mo agoHugging Face19altaidevorg /fineweb2-hq-turkishtext1M<n<10M0 likes554 downloads11mo agoHugging Face20ytu-ce-cosmos /Cosmos-Turkish-Corpus-v1.0This is the Turkish pretraining corpus of the Cosmos AI Research Group. It contains ~15B tokens and demonstrates competitive performance across various Turkish benchmarks when used in continual pretraining setups. Cosmos-Turkish-Corpus is collected from a wide range of Turkish websites, including forums, news sources, Wikipedia, and more. URL-based deduplication has been applied; however, additional content-level deduplication and filtering may be required before use. text1M<n<10M26 likes539 downloads10mo agoHugging Face213nesdeniz /turkish-daily-dialogues-5k Turkish Daily Dialogues 5K Exactly 5,000 synthetic, multi-turn Turkish conversations covering ordinary daily-life situations. The corpus is designed as a small, auditable baseline for dialogue modelling, instruction-format experiments, augmentation research, and Turkish-language evaluation—not as a substitute for conversations written by real people. Provenance in one sentence: the Turkish source scenario library was drafted with AI assistance specifically for this project… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/turkish-daily-dialogues-5k.texttext-generation1K<n<10K2 likes510 downloads2mo agoHugging Face22ituperceptron /image-captioning-turkish Türkçe Image Captioning Veri Seti Bu veri seti BLIP3o modelinin pretrain eğitiminde kullanılan BLIP3o-Pretrain-Long-Caption ve BLIP3o-Pretrain-Short-Caption veri setlerinin Türkçeye çevirilmiş bir alt parçasıdır. Orijinal veri setinin oluşturulması ile ilgili detaylı bilgiye BLIP-3o makalesi üzerinden ulaşabilirsiniz. Veri seti Image-to-Text modellerinin eğitilmesinde veya ince ayar sürecinde kullanılabilir. Veri seti, orijinal veri setinin lisansı olan Apache 2.0 altında… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/image-captioning-turkish.imageimage-to-text1M<n<10M7 likes504 downloads8mo agoHugging Face23umarigan /turkish_corpustextfeature-extraction10M<n<100M5 likes489 downloads3y agoHugging Face24AdnanElAssadi /Google-Translated_Turkish_GPQA_Datasettabular1K<n<10K0 likes485 downloads2y agoHugging Face25selmanbaysan /cleaned_turkish_embedding_model_training_data_colabtext10M<n<100M1 likes483 downloads1y agoHugging Face26turkish-nlp-suite /temiz-OSCAR Dataset Card for Temiz OSCAR Temiz OSCAR is a corpora collection consisting of cleaned versions of original OSCAR corpora. This collection is made up of four datasets: OSCAR-2019, OSCAR-2109, OSCAR-2201 and OSCAR-2301 This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. Dataset num instances size num of words OSCAR-2019 3.671.430 7.7G 976M OSCAR-2109 8.472.809 18G 2.22B OSCAR-2201… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-OSCAR.textfill-mask10M<n<100M5 likes479 downloads11mo agoHugging Face27Anilosan15 /Turkish_TTS_Dataaudiotext-to-speech10K<n<100K21 likes448 downloads7mo agoHugging Face28AlicanKiraz0 /Turkish-SFT-Dataset-v1.0 Turkish-SFT-Dataset-v1.01 Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci 🔎 Özet Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.texttext-classification1K<n<10K52 likes437 downloads11mo agoHugging Face29TFLai /Turkish-AlpacaStanford alpaca turkish: Stanford Alpaca text10K<n<100K28 likes434 downloads3y agoHugging Face30trmteb /cleaned_turkish_embedding_model_training_data_colab Citation If you use this dataset in your research, please cite the following paper: @inproceedings{baysan-gungor-2025-tr, title = "{TR}-{MTEB}: A Comprehensive Benchmark and Embedding Model Suite for {T}urkish Sentence Representations", author = "Baysan, Mehmet Selman and Gungor, Tunga", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2025", month = nov, year = "2025", address = "Suzhou, China", publisher =… See the full description on the dataset page: https://huggingface.co/datasets/trmteb/cleaned_turkish_embedding_model_training_data_colab.text10M<n<100M3 likes423 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.