CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01altaidevorg /fineweb-2-turkish-categorized What is this THis is the categorized version of the Turkish subset of the fineweb-2 dataset. It is an ongoing effort, and the details will be added soon with the rest of the dataset. tabular10M<n<100M15 likes3.5k downloads2y agoHugging Face02moganai /turkishfineweb2-cleaned TurkishFineweb2-Cleaned A Turkish web corpus derived from the Turkish (tur_Latn) subset of FineWeb-2, augmented with an additional quality-classification layer and a near-duplicate removal pass. 📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum Source FineWeb-2 is a large-scale, multilingual web corpus built from Common Crawl. This dataset covers the Turkish (tur_Latn) portion of FineWeb-2, spanning the… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkishfineweb2-cleaned.tabulartext-generation10M<n<100M5 likes1.7k downloads2d agoHugging Face03Ethosoft /Turkish_corpus Turkish Corpus 🇹🇷 Turkish Corpus is a large-scale cleaned Turkish text dataset created by collecting public Turkish corpora and extracting Turkish-language portions from multilingual datasets. The dataset is designed for Turkish Natural Language Processing research, language model pretraining, tokenizer training, embedding models, retrieval systems, and general Turkish language understanding tasks. The main purpose of this dataset is to provide a practical, scalable, and… See the full description on the dataset page: https://huggingface.co/datasets/Ethosoft/Turkish_corpus.tabulartext-generation1M<n<10M1 likes1.4k downloads5mo agoHugging Face04mrfg /turkish-court-decisions Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/mrfg/turkish-court-decisions.tabulartext-generation10M<n<100M5 likes1.3k downloads1mo agoHugging Face05selmanbaysan /turkish_embedding_model_training_datatextsentence-similarity100M<n<1B5 likes1.3k downloads1y agoHugging Face06moganai /mogan-turkish-web Mogan Turkish Web A large-scale Turkish web corpus derived from monthly Common Crawl snapshots covering the period from January 2025 to June 2026. The corpus was constructed by extracting Turkish-language content from raw Common Crawl WARC/WET dumps, followed by language filtering, PII masking, and near-duplicate removal. 📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum Dataset Summary This dataset… See the full description on the dataset page: https://huggingface.co/datasets/moganai/mogan-turkish-web.texttext-generation10M<n<100M6 likes1.2k downloads2d agoHugging Face07hasankursun /turkish-corpus-100b Turkish Corpus 100B (TC-100B) Dataset Summary The Turkish Corpus 100B (TC-100B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Turkish. Comprising approximately 105 Billion tokens (measured with Qwen/Llama3 tokenizer), it represents one of the largest open resources for Turkish LLM pretraining. The dataset is engineered for a two-stage training pipeline: Pretrain Subset (~103B Tokens): A diverse mix of… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/turkish-corpus-100b.texttext-generation100M<n<1B8 likes846 downloads3mo agoHugging Face08Taklaxbr /turkish-tts-combined-raw Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından düzenlenmiştir. Orijinal veri seti afkfatih tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir. 🔗 Orijinal Kaynak: afkfatih/turkish-tts-combined-raw 🔗 Derleyen Platform: VeriPazarı Türkçe TTS Birleşik Veri Seti (Turkish TTS Combined) 7 farklı açık kaynak Türkçe TTS (Metinden Sese) veri setinin birleşimidir. ~81.500 örnek |… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/turkish-tts-combined-raw.audiotext-to-speech10K<n<100K0 likes791 downloads4mo agoHugging Face09altaidevorg /fineweb2-hq-turkishtext1M<n<10M0 likes554 downloads11mo agoHugging Face10ytu-ce-cosmos /Cosmos-Turkish-Corpus-v1.0This is the Turkish pretraining corpus of the Cosmos AI Research Group. It contains ~15B tokens and demonstrates competitive performance across various Turkish benchmarks when used in continual pretraining setups. Cosmos-Turkish-Corpus is collected from a wide range of Turkish websites, including forums, news sources, Wikipedia, and more. URL-based deduplication has been applied; however, additional content-level deduplication and filtering may be required before use. text1M<n<10M26 likes539 downloads10mo agoHugging Face113nesdeniz /turkish-daily-dialogues-5k Turkish Daily Dialogues 5K Exactly 5,000 synthetic, multi-turn Turkish conversations covering ordinary daily-life situations. The corpus is designed as a small, auditable baseline for dialogue modelling, instruction-format experiments, augmentation research, and Turkish-language evaluation—not as a substitute for conversations written by real people. Provenance in one sentence: the Turkish source scenario library was drafted with AI assistance specifically for this project… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/turkish-daily-dialogues-5k.texttext-generation1K<n<10K2 likes510 downloads2mo agoHugging Face12ituperceptron /image-captioning-turkish Türkçe Image Captioning Veri Seti Bu veri seti BLIP3o modelinin pretrain eğitiminde kullanılan BLIP3o-Pretrain-Long-Caption ve BLIP3o-Pretrain-Short-Caption veri setlerinin Türkçeye çevirilmiş bir alt parçasıdır. Orijinal veri setinin oluşturulması ile ilgili detaylı bilgiye BLIP-3o makalesi üzerinden ulaşabilirsiniz. Veri seti Image-to-Text modellerinin eğitilmesinde veya ince ayar sürecinde kullanılabilir. Veri seti, orijinal veri setinin lisansı olan Apache 2.0 altında… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/image-captioning-turkish.imageimage-to-text1M<n<10M7 likes504 downloads8mo agoHugging Face13umarigan /turkish_corpustextfeature-extraction10M<n<100M5 likes489 downloads3y agoHugging Face14AdnanElAssadi /Google-Translated_Turkish_GPQA_Datasettabular1K<n<10K0 likes485 downloads2y agoHugging Face15selmanbaysan /cleaned_turkish_embedding_model_training_data_colabtext10M<n<100M1 likes483 downloads1y agoHugging Face16Anilosan15 /Turkish_TTS_Dataaudiotext-to-speech10K<n<100K21 likes448 downloads7mo agoHugging Face17trmteb /cleaned_turkish_embedding_model_training_data_colab Citation If you use this dataset in your research, please cite the following paper: @inproceedings{baysan-gungor-2025-tr, title = "{TR}-{MTEB}: A Comprehensive Benchmark and Embedding Model Suite for {T}urkish Sentence Representations", author = "Baysan, Mehmet Selman and Gungor, Tunga", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2025", month = nov, year = "2025", address = "Suzhou, China", publisher =… See the full description on the dataset page: https://huggingface.co/datasets/trmteb/cleaned_turkish_embedding_model_training_data_colab.text10M<n<100M3 likes423 downloads10mo agoHugging Face18umarigan /turkish_corpus_tokenized Dataset Card for "turkish_corpus_tokenized" More Information needed 10M<n<100M0 likes409 downloads3y agoHugging Face19projectkaira /turkish-tts-combined-raw Türkçe TTS Birleşik Veri Seti 7 farklı açık kaynak Türkçe TTS veri setinin birleşimi. ~81,500 örnek | 24kHz | SNAC uyumlu Kaynaklar Veri Seti Örnek Kaynak Mazlum Kiper 9,643 omersaidd/tts_mazlum_kiper_tur Ahmet Deniz 11,289 omersaidd/tts_ahmet_deniz_tur Nisan Kumru 8,042 omersaidd/tts_nisan_kumru_tur Derya TTS v2 42 afkfatih/derya-tts-v2 Derya Karma v3 255 afkfatih/derya-tts-karma-v3 Khan Academy 25,741 ysdede/khanacademy-turkish Common… See the full description on the dataset page: https://huggingface.co/datasets/projectkaira/turkish-tts-combined-raw.audiotext-to-speech10K<n<100K1 likes405 downloads2mo agoHugging Face20boun-tabilab /turkish_parliamentary_data Grand National Assembly Corpus of Türkiye (GNACT) A comprehensive collection of Turkish parliamentary transcripts spanning over 100 years (1920–present), from 10 legislative bodies. Includes both Ottoman Turkish (1920–1928) and Modern Turkish (1928–present) texts. Loading the dataset from datasets import load_dataset # Strategy 1: full session documents, all bodies (default) ds = load_dataset("boun-tabilab/turkish_parliamentary_data", "full_sessions", split="train") #… See the full description on the dataset page: https://huggingface.co/datasets/boun-tabilab/turkish_parliamentary_data.tabulartext-generation1M<n<10M8 likes397 downloads6mo agoHugging Face21afkfatih /turkish-tts-combined-raw Türkçe TTS Birleşik Veri Seti 7 farklı açık kaynak Türkçe TTS veri setinin birleşimi. ~81,500 örnek | 24kHz | SNAC uyumlu Kaynaklar Veri Seti Örnek Kaynak Mazlum Kiper 9,643 omersaidd/tts_mazlum_kiper_tur Ahmet Deniz 11,289 omersaidd/tts_ahmet_deniz_tur Nisan Kumru 8,042 omersaidd/tts_nisan_kumru_tur Derya TTS v2 42 afkfatih/derya-tts-v2 Derya Karma v3 255 afkfatih/derya-tts-karma-v3 Khan Academy 25,741 ysdede/khanacademy-turkish Common Voice 17 26… See the full description on the dataset page: https://huggingface.co/datasets/afkfatih/turkish-tts-combined-raw.audiotext-to-speech10K<n<100K15 likes381 downloads10mo agoHugging Face22umarigan /turkish_clip_dataset_with_text_embeddingsThis dataset cleaned and dowloaded version of following dataset: https://huggingface.co/datasets/visheratin/laion-coco-nllb The main purpose was to extract Turkish captions and download images. You can use this dataset to fine-tune or create a clip model. Since there English and Turkish captions you can also use those to create language model? image100K<n<1M1 likes380 downloads3y agoHugging Face23serdarcaglar /turkish-tts-audiobooksgated Turkish TTS Audiobooks Turkish read-speech corpus for text-to-speech training, built from Turkish audiobook and spoken-article recordings by an automatic pipeline: VAD segmentation → technical QC → acoustic event tagging → DNSMOS → speaker embedding/consistency → double-pass Whisper ASR → text policy → leakage-free splitting. Audio is 16 kHz mono lossless FLAC embedded in the Parquet shards. The pipeline that produced it — every stage, every threshold, the export and audit… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/turkish-tts-audiobooks.audiotext-to-speech100K<n<1M9 likes380 downloads1mo agoHugging Face24toksuite /toksuite_turkish Dataset Card for Tokenization Robustness TokSuite Benchmark (Turkish Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Turkish language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness. Curated by: R3 Research Team… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_turkish.textmultiple-choicen<1K0 likes342 downloads8mo agoHugging Face25GoktugD /turkish-nli-constructed-1.5m Turkish NLI Constructed 1.5M v2 Dengeli entailment, contradiction ve neutral sınıflı Türkçe NLI çiftleri. Doğrulanmış boyut Train: 1,470,000 Validation: 15,000 Test: 15,000 Toplam: 1,500,000 Ana görev sütunları: id, premise, hypothesis, label Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-nli-constructed-1.5m.texttext-classification1M<n<10M0 likes326 downloads2mo agoHugging Face26Alptekinege /turkish-court-decisions Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Alptekinege/turkish-court-decisions.tabulartext-generation10M<n<100M3 likes326 downloads24d agoHugging Face27erenfazlioglu /turkishvoicedataset Dataset Card for "turkishneuralvoice" Dataset Overview Dataset Name: Turkish Neural Voice Description: This dataset contains Turkish audio samples generated using Microsoft Text to Speech services. The dataset includes audio files and their corresponding transcriptions. Dataset Structure Configs: default Data Files: Split: train Path: data/train-* Dataset Info: Features: audio: Audio file transcription: Corresponding text transcription Splits: train… See the full description on the dataset page: https://huggingface.co/datasets/erenfazlioglu/turkishvoicedataset.audiotext-to-speech100K<n<1M44 likes318 downloads2y agoHugging Face28ituperceptron /image-vqa-turkish Türkçe Image VQA Veri Seti Bu veri seti Türkçe görsel soru-cevap (VQA) çiftleri içermektedir. Kullanım from datasets import load_dataset ds = load_dataset("ituperceptron/turkish-image-vqa", split="vqa") Veri Yapısı image: Görsel (PIL Image) image_id: Görselin benzersiz ID'si vqa: VQA soru-cevap çiftleri (JSON formatında) imagevisual-question-answering100K<n<1M3 likes307 downloads9mo agoHugging Face29fthbrmnby /turkish_product_reviews Dataset Card for Turkish Product Reviews Dataset Summary This Turkish Product Reviews Dataset contains 235.165 product reviews collected online. There are 220.284 positive, 14881 negative reviews. Supported Tasks and Leaderboards [More Information Needed] Languages The dataset is based on Turkish. Dataset Structure Data Instances Example 1: sentence: beklentimin altında bir ürün kaliteli değil sentiment: 0 (negative) Example 2:… See the full description on the dataset page: https://huggingface.co/datasets/fthbrmnby/turkish_product_reviews.texttext-classification100K<n<1M15 likes299 downloads2y agoHugging Face30akuzdeuov /turkish_maleaudio10K<n<100K0 likes276 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.