CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01turkish-nlp-suite /BellaTurca Dataset Card for BellaTurca BellaTurca is the first large-scale Turkish corpus collection for training Turkish language models. The total size is around 245GB and 30 billion words. BellaTurca's focus is high quality, diversity as well as the size. This collection is made up of five datasets: AkademikDerlem, OzenliDerlem, ForumSohbetleri, Temiz OSCAR and Temiz mC4. Originally there was a book corpus included, but it is excluded due to containing copyrighted material. AkademikDerlem… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/BellaTurca.text10M<n<100M17 likes3.5k downloads7mo agoHugging Face02turkish-nlp-suite /OzenliDerlem Dataset Card for OzenliDerlem OzenliDerlem (a.k.a CraftedCrawl) is a carefully assembled collection of top-notch web crawl data from handpicked websites, featuring articles, journals, and magazines. It focuses on gathering rich and detailed text content, especially longer articles. The collection covers a wide range of topics, including travel, news, culture, fairy tales and folklore, movie reviews, popular science, product and service complaints, fashion and self-care, trendy… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/OzenliDerlem.text1M<n<10M11 likes712 downloads7mo agoHugging Face03turkish-nlp-suite /TrGLUE TrGLUE - The First Non-Translate Natural Language Understanding Benchmark for Turkish Dataset Card for TrGLUE TrGLUE is a natural language understanding benchmarking dataset including several single sentence and sentence pair classification tasks. The inspiration is clearly the original GLUE benchmark. Tasks Single Sentence Tasks TrCOLA The original Corpus of Linguistic Acceptability consists of sentences compiled from English literature textbooks.… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/TrGLUE.texttext-classification100K<n<1M6 likes567 downloads9mo agoHugging Face04turkish-nlp-suite /temiz-OSCAR Dataset Card for Temiz OSCAR Temiz OSCAR is a corpora collection consisting of cleaned versions of original OSCAR corpora. This collection is made up of four datasets: OSCAR-2019, OSCAR-2109, OSCAR-2201 and OSCAR-2301 This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. Dataset num instances size num of words OSCAR-2019 3.671.430 7.7G 976M OSCAR-2109 8.472.809 18G 2.22B OSCAR-2201… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-OSCAR.textfill-mask10M<n<100M5 likes479 downloads11mo agoHugging Face05turkish-nlp-suite /AkademikDerlem Dataset Card for AkademikDerlem AkademikDerlem is a scientific text corpus for Turkish, gathered from misc academical publication websites. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. This collection is made up of five datasets: Articles, Academic-Abstracts, Medical-Articles, Medical-Abstracts, and Bilkent-Writings. The Bilkent-Writings dataset comes from creative writings produced in the… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/AkademikDerlem.textfill-mask100K<n<1M6 likes250 downloads11mo agoHugging Face06turkish-nlp-suite /Havadis Dataset Card for Havadis Havadis is a high quality and large Turkish news corpus, indeed the largest Turkish news corpus ever. This corpus is scraped from online news sebsites and includes text from popular newspapers such as CNN Türk Habertürk Hürriyet Millyet NTV Posta Sabah Star Sözcü Takvim . The instances are first crawled from the corresponding websites, then went throught an extensive cleaning process. We eliminated instances that are too short, too repetetive… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Havadis.textfill-mask100K<n<1M6 likes233 downloads2mo agoHugging Face07turkish-nlp-suite /TurkishHateMap Turkish Hate Map - A Large Scale and Diverse Hate Speech Dataset for Turkish Dataset Summary Turkish Hate Map (TuHaMa for short) is a big scale Turkish hate speech dataset that includes diverse target groups such as misogyny, political animosity, animal aversion, vegan antipathy, ethnic group hostility, and more. The dataset includes a total of 52K instances with 13 target groups. The dataset includes 4 labels, offensive, hate, neutral and civilized. Here is the… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/TurkishHateMap.texttext-classification10K<n<100K4 likes212 downloads2y agoHugging Face08turkish-nlp-suite /InstrucTurca InstrucTurca v1.0.0 is a diverse synthetic instruction tuning dataset crafted for instruction-tuning Turkish LLMs. The data is compiled data various English datasets and sources, such as code instructions, poems, summarized texts, medical texts, and more. Dataset content BI55/MedText checkai/instruction-poems garage-bAInd/Open-Platypus Locutusque/ColumnedChatCombined nampdn-ai/tiny-codes Open-Orca/OpenOrca pubmed_qa TIGER-Lab/MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/InstrucTurca.texttext-generation1M<n<10M40 likes204 downloads2y agoHugging Face09turkish-nlp-suite /ForumSohbetleri Dataset Card for ForumSohbetleri ForumSohbetleri a web forum tetx corpus for Turkish, indeed first large-scale Turkish forum text corpus. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. This collection is made up of several subsets, each subset is gathered from the corresponding forum website. Forum websites contains diverse topics, ladies only, tech, economics, life, relations and much more...… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/ForumSohbetleri.textfill-mask1M<n<10M5 likes201 downloads11mo agoHugging Face10turkish-nlp-suite /turkish-wikiNER Dataset Card for "turkish-nlp-suite/turkish-wikiNER" Dataset Summary Turkish NER dataset from Wikipedia sentences. 20.000 sentences are sampled and re-annotated from Kuzgunlar NER dataset. Annotations are done by Co-one. Many thanks to them for their contributions. This dataset is also used in our brand new spaCy Turkish packages. Dataset Instances An instance of this dataset looks as follows: { "tokens": ["Çekimler", "5", "Temmuz", "2005", "tarihinde"… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/turkish-wikiNER.texttoken-classification10K<n<100K14 likes182 downloads5mo agoHugging Face11turkish-nlp-datasets /eksisozluk-ekonomi-ve-finans-tr Ekşi Sözlük Türkçe Teknoloji Dataset Ekşi Sözlük'ten derlenen, teknoloji kategorisine ait Türkçe kullanıcı entry'lerinden oluşan bir veri setidir. Türkçe NLP araştırmaları ve LLM eğitimi için hazırlanmıştır. İçerik Makroekonomik kavramlar, yatırım araçları, kripto para ve güncel ekonomik gelişmeleri kapsayan 39 farklı başlık altında toplanmış entry'lerden oluşmaktadır. Alan Kapsanan Başlıklar Ekonomik Kriz & Genel Durum 2025-2026 ekonomik krizi, devalüasyon… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-datasets/eksisozluk-ekonomi-ve-finans-tr.text10K<n<100K0 likes144 downloads5mo agoHugging Face12turkish-nlp-suite /SentiTurca SentiTurca - A Sentiment Analysis Benchmark for Turkish Dataset Card for SentiTurca SentiTurca is a sentiment analysis benchmarking dataset including movie reviews, hate speech and e-commerce reviews classification. Datasets e-commerce: The e-commerce reviews are scraped from e-commerce websites Trendyol.com and Hepsiburada.com, including review for many product types such as cloths, toys, books, electronics and more.E-commerce reviews has their stand… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/SentiTurca.texttext-classification100K<n<1M7 likes96 downloads5mo agoHugging Face13turkish-nlp-suite /turkish-morph-analysis Dataset Card for TrMorphTester This dataset is a testing dataset for Turkish morphology, aiming to calculate how other subword strategies aligns with morphological segmentation of Turkish. The data is automatically generated from Turkish morpoholigical lexicon. For each row, we offer a surface form, then lemma and all suffixes, separated by a + character. The dataset has several splits for several purposes: lemma: Surface form is same with lemma, no suffixes at all. Nouns, verbs… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/turkish-morph-analysis.text100K<n<1M2 likes77 downloads7mo agoHugging Face14turkish-nlp-suite /vitamins-supplements-reviews Dataset Card for turkish-nlp-suite/vitamins-supplements-reviews Dataset Summary Turkish sentiment analysis dataset from customer reviews about supplement and vitamin products. The dataset is scraped from Vitaminler.com and contains customer reviews and star rating about vitamin and supplement products. Each customer review in the Vitamins and Supplements Reviews Dataset describes a customer’s experience with a supplement product in terms of the product’s effectiveness… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/vitamins-supplements-reviews.texttext-classification100K<n<1M1 likes73 downloads2y agoHugging Face15turkish-nlp-suite /temiz-WikiA cleaned version of Turkish Wikipedia dataset. The soource is wikimedia/wikipedia repo. The text is cleaned throughoutly, first of all we eliminated text that are shorter than a predetermined threshold of words and characters. Then we normalized with NFKC, cleaned some non-ASCII chars and normalized whitespaces. textfill-mask100K<n<1M4 likes64 downloads7mo agoHugging Face16turkish-nlp-suite /beyazperde-top-300-movie-reviews Dataset Card for turkish-nlp-suite/beyazperde-top-300-movie-reviews Dataset Summary Beyazperde Movie Reviews offers Turkish sentiment analysis datasets that is scraped from popular movie reviews website Beyazperde.com. Top 300 Movies include audience reviews about best 300 movies of all the time. Here's the star rating distribution: star rating count 0.5 101 1.0 39 1.5 19 2.0 44 2.5 210 3.0 196 3.5 490 4.0 1212 4.5 818 5.0 1251 total 4380… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/beyazperde-top-300-movie-reviews.texttext-classification1K<n<10K2 likes58 downloads2y agoHugging Face17turkish-nlp-suite /beyazperde-all-movie-reviews Dataset Card for turkish-nlp-suite/beyazperde-all-movie-reviews Dataset Summary Beyazperde Movie Reviews offers Turkish sentiment analysis datasets that is scraped from popular movie reviews website Beyazperde.com. All Movie Reviews include audience reviews about movies of all the time. Here's the star rating distribution: star rating count 0.5 3.635 1.0 2.325 1.5 1.077 2.0 1.902 2.5 4.767 3.0 4.347 3.5 6.495 4.0 9.486 4.5 3.652 5.0 7.594… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/beyazperde-all-movie-reviews.texttext-classification10K<n<100K1 likes46 downloads2y agoHugging Face18turkish-nlp-suite /MusteriYorumlari MüşteriYorumlari - A Large Scale Customer Sentiment Analysis Dataset for Turkish Dataset Summary MüşteriYorumları is a Turkish e-commerce customer reviews dataset of size 103K, scraped from Hepsiburada.com and Trendyol.com. These reviews encompass a wide array of product categories, including apparel, food items, baby products, and books. Review stars are in range of 1-5 stars. The star distribution is as follows: star rating count 1 12,873 2 11,472 3 18… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/MusteriYorumlari.texttext-classification100K<n<1M3 likes44 downloads2y agoHugging Face19turkish-nlp-datasets /economy-and-finance Ekşi Sözlük Türkçe Teknoloji Dataset Ekşi Sözlük'ten derlenen, teknoloji kategorisine ait Türkçe kullanıcı entry'lerinden oluşan bir veri setidir. Türkçe NLP araştırmaları ve LLM eğitimi için hazırlanmıştır. İçerik Yapay zeka, ChatGPT, sosyal medya algoritmaları, kripto para ve diğer teknoloji konularını kapsayan 24 farklı başlık altında toplanmış entry'lerden oluşmaktadır. Alan Kapsanan Başlıklar Yapay Zeka yapay zeka, chatgpt, claude ai, meta ai, apple… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-datasets/economy-and-finance.text1K<n<10K0 likes40 downloads5mo agoHugging Face20turkish-nlp-datasets /eksisozluk_teknoloji_tr Ekşi Sözlük Türkçe Teknoloji Dataset Genel Bilgi Dil: Türkçe Kaynak: Ekşi Sözlük Konu: Teknoloji (Yapay Zeka, ChatGPT, Sosyal Medya vb.) Format: JSONL Toplam Entry: 7521 Alan Açıklamaları text: Entry içeriği author: Yazar nickname date: Giriş tarihi entry_no: Tekil entry ID topic: Başlık kategorisi source: eksisozluk language: tr Kullanım Amacı Türkçe LLM eğitimi ve fine-tuning için uygundur. Lisans CC BY-NC 4.0 text1K<n<10K0 likes33 downloads5mo agoHugging Face21turkish-nlp-suite /Corona-mini Dataset Card for turkish-nlp-suite/Corona-mini Dataset Summary This is a tiny Turkish corpus consisting of comments about Corona symptoms. The corpus is compiled from two Ekşisözlük headlines "covid-19 belirtileri" and "gün gün koronavirüs belirtileri": https://eksisozluk.com/covid-19-belirtileri--6416646 https://eksisozluk.com/gun-gun-koronavirus-belirtileri--6757665 This corpus contains 178 raw, 175 processed comments all comments are in Turkish comes in 2… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Corona-mini.textsummarizationn<1K0 likes30 downloads3y agoHugging Face22turkish-nlp-suite /sinefil-movie-reviews Sinefil Movie Reviews A movie reviews sentiment analysis dataset for Turkish. Dataset Summary Sinefil Movie Reviews offers Turkish sentiment analysis datasets that is scraped from popular movie reviews website Sinefil.com. The reviews include audience reviews about movies of all the time. The score field takes values between 1 and 9.9. Values are like 8, 8.1, 8.2 .. 8.9. Here's the distribution divided into integer bins: star rating count 1-2 2323 2-3 874… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/sinefil-movie-reviews.texttext-classification10K<n<100K0 likes24 downloads2y agoHugging Face23turkish-nlp-suite /TrCOLA TrCOLA - Corpus of Linguistic Acceptability for Turkish Language Dataset Card for TrCOLA TrCOLA is the Turkish version of CoLA dataset, The Corpus of Linguistic Acceptability. This dataset introduces linguistic acceptability task for Turkish. The total dataset size is 9.9K instances. Each instance of the dataset is an original and correct sentence, variation of sentence that is produced in a specific way, the variation type and a binary label stating the sentence is a… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/TrCOLA.texttext-classification1K<n<10K2 likes23 downloads2y agoHugging Face24turkish-nlp-suite /BuyukSinema BüyükSinema - A Large Scale Turkish Movie Reviews Sentiment Dataset Dataset Summary BüyükSinema is a Turkish movie reviews dataset of size 87K, scraped from Sinefil.com and Beyazperde.com. Hence this dataset is a superset of BeyazPerde All Movie Reviews, BeyazPerde Top 300 Movie Reviews and Sinefil Movie Reviews datasets. This is a merge of the three different datasets from two resources, hence we scaled the output stars into the range of 1-10 accordingly. The star… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/BuyukSinema.texttext-classification10K<n<100K1 likes23 downloads2y agoHugging Face25turkish-nlp-suite /vitamins-supplements-NER Dataset Card for turkish-nlp-suite/vitamins-supplements-NER Dataset Summary The Vitamins and Supplements NER Dataset is a NER dataset containing customer reviews with entity and span annotations. User reviews were collected from a popular supplement products e- commerce website Vitaminler.com. Each customer review in the Vitamins and Supplements NER Dataset describes a customer’s experience with a supplement product in terms of that product’s effectiveness, side… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/vitamins-supplements-NER.texttoken-classification1K<n<10K2 likes21 downloads3y agoHugging Face26turkish-nlp-suite /Treebank-Benchmarking Turkish Treebank Benchmarking This is the repo for Turkish treebank benchmarking, namely evaluating Tranformer models on POS-Dep-Morph task. For the data, we used two treebank, IMST and BOUN. We converted conllu format to json lines for being compatible to HF dataset formats. Here are treebank sizes at a glance: Dataset train lines dev lines test lines BOUN 7803 979 979 IMST 3435 1100 1100 A typical instance from the dataset looks like: { "id": "ins_1267"… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Treebank-Benchmarking.text10K<n<100K0 likes21 downloads7mo agoHugging Face27STNM-NLPhoenix /turkish-org-ner Turkish Organization Named Entity Recognition (NER) Dataset This dataset is designed for Named Entity Recognition (NER) tasks in the Turkish language, specifically focusing on organization entities. It contains sentences annotated with the ORGANIZATION label, making it a valuable resource for training and evaluating NER models. Usage To use this dataset, you can load it directly from Huggingface's datasets library: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/STNM-NLPhoenix/turkish-org-ner.text1M<n<10M1 likes6 downloads2y agoHugging Face28STNM-NLPhoenix /turkish-absa Turkish ABSA Dataset This repository contains a dataset created for training and evaluating an Aspect-Based Sentiment Analysis (ABSA) model in the Turkish language. The dataset includes sentences with various entities and sentiments. About the Dataset The dataset consists of sentences taken from customer feedback and includes tags related to the entities and sentiments mentioned in these sentences. Each review indicates which entity it is related to and the sentiment… See the full description on the dataset page: https://huggingface.co/datasets/STNM-NLPhoenix/turkish-absa.text100K<n<1M4 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.