datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish_instructionsturkish-sentiment-analysis-dataset
Dataset
This dataset contains positive , negative and notr sentences from several data sources given in the references. In the most sentiment models , there are only two labels; positive and negative. However , user input can be totally notr sentence. For such cases there were no data I could find. Therefore I created this dataset with 3 class. Positive and negative sentences are listed below. Notr examples are extraced from turkish wiki dump. In addition, added some random text… See the full description on the dataset page: https://huggingface.co/datasets/winvoker/turkish-sentiment-analysis-dataset.Flutter-Code-with-Questions-Dataset-Turkish
Flutter Code with Questions Dataset (Turkish)
📦 Dataset Name: flutter_code_with_questions
Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir.
📁 Dataset Format
Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.stsb-mt-turkish
STSb Turkish
Semantic textual similarity dataset for the Turkish language. It is a machine translation (Azure) of the STSb English dataset. This dataset is not reviewed by expert human translators.
Uploaded from this repository.
Citing & Authors
@misc{celik2020stsbtr,
author = {Emrecan Çelik},
title = {STSB-MT-Turkish},
howpublished = {Hugging Face dataset repository},
url = {https://huggingface.co/datasets/emrecan/stsb-mt-turkish}… See the full description on the dataset page: https://huggingface.co/datasets/emrecan/stsb-mt-turkish.turkish-law-corpus
⚖️ Turkish Law — 106 Kanun Korpusu & Soru-Cevap106 Statutes Corpus & QA
🇹🇷 Türk hukukunun en çok kullanılan 106 kanunu, madde madde temizlenmiş 16.001 metin parçası ve bu maddelere dayalı 5.011 Türkçe soru-cevap çifti. Tamamı resmî kaynaktan (mevzuat.gov.tr), RAG ve yapay zekâ uygulamaları için hazır.
🇬🇧 The 106 most widely used Turkish statutes as 16,001 clean, article-level text chunks, plus 5,011 Turkish question-answer pairs grounded in those articles. All from the… See the full description on the dataset page: https://huggingface.co/datasets/CtnkyaABC/turkish-law-corpus.turkish-toxic-language
Turkish Texts for Toxic Language Detection
Dataset Description
Dataset Summary
This text dataset is a collection of Turkish texts that have been merged from various existing offensive language datasets found online. The dataset contains a total of 77,800 instances, each labeled as either offensive or not offensive.
To ensure the dataset's completeness, we utilized multiple transformer models to augment the dataset with pseudo labels. The resulting dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Overfit-GM/turkish-toxic-language.turkish-offensive-language-detection
Dataset Summary
This dataset is enhanced version of existing offensive language studies. Existing studies are highly imbalanced, and solving this problem is too costly. To solve this, we proposed contextual data mining method for dataset augmentation. Our method is basically prevent us from retrieving random tweets and label individually. We can directly access almost exact hate related tweets and label them directly without any further human interaction in order to solve imbalanced… See the full description on the dataset page: https://huggingface.co/datasets/Toygar/turkish-offensive-language-detection.turkish-hate-speech-superset
Turkish Hate Speech Superset
This dataset is a superset (N=41,423) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Turkish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/turkish-hate-speech-superset.Turkish_SentimentAnalysis_TRSAv1TRSAv1 (Turkish Sentiment Analysis Version 1) Dataset
This data set has been produced to contribute to Turkish NLP studies.
The dataset consists of a total of 150 thousand samples, 50 thousand negative, 50 thousand positive, and 50 thousand neutral.
It can be used in text classification and sentiment analysis studies by citing the related study.
Related Work
Aydoğan M, Kocaman V. TRSAv1: A new benchmark dataset for classifying user reviews on Turkish e-commerce websites. Journal of… See the full description on the dataset page: https://huggingface.co/datasets/maydogan/Turkish_SentimentAnalysis_TRSAv1.turkish-disaster-news-geonlp
Turkish Disaster News GeoNLP Dataset
Dataset Summary
This dataset was created and submitted as part of the Uncharted Data Challenge by Adaption. The LLM-enhanced instruction pairs (turkish_earthquake_news.csv) were generated using Adaptive Data by Adaption — an AI-powered data adaptation platform.
The first open-source Turkish-language disaster news dataset with district-level geocoding, humanitarian category labels, and multi-dimensional damage classification.… See the full description on the dataset page: https://huggingface.co/datasets/FatmaElik/turkish-disaster-news-geonlp.turkish-plu-goal-inferenceHomepage: https://github.com/GGLAB-KU/turkish-plu
turkish-google-maps-15M
Turkish Google Maps Reviews
Bu veri seti, Türkiye’deki işletmelere ait Türkçe Google Maps yorumlarını içerir.
Her kayıt:
yorum metni
yorum puanı
işletme adı
işletme kategorisi
gibi bilgileri içerir.
Veri seti, özellikle büyük ölçekli Türkçe NLP çalışmaları için uygundur.
Contents
Veri setinde aşağıdaki türde alanlar bulunmaktadır:
yorum metni (review_text)
yorum puanı (rating)
işletme adı (place_name)
işletme kategorisi (category)
kategori listesi (category_list)… See the full description on the dataset page: https://huggingface.co/datasets/opdullah/turkish-google-maps-15M.TARA_Turkish_LLM_Benchmark
TARA: Turkish Advanced Reasoning Assessment Veri Seti
*Img Credit: Open AI ChatGPT
**English version is given below.**
Evaluation Notebook / Değerlendirme Not Defteri
Dataset Summary
TARA (Turkish Advanced Reasoning Assessment), Türkçe dilindeki Büyük Dil Modellerinin (LLM'ler) gelişmiş akıl yürütme yeteneklerini çoklu alanlarda ölçmek için tasarlanmış, zorluk derecesine göre sınıflandırılmış bir benchmark veri setidir. Bu veri seti, LLM'lerin sadece bilgi… See the full description on the dataset page: https://huggingface.co/datasets/emre/TARA_Turkish_LLM_Benchmark.turkish_news_datasetturkish-plu-step-inferenceHomepage: https://github.com/GGLAB-KU/turkish-plu
turkish-plu-step-orderingHomepage: https://github.com/GGLAB-KU/turkish-plu/
turkish-legal-rag
Turkish Legal RAG Corpus — Türk Hukuku için Açık RAG Datasetı
Tek cümle: 25 önemli Türk kanununun (mevzuat.gov.tr kaynaklı, madde bazlı temiz chunk'lar) + 290 manuel doğrulanmış soru-cevap altın benchmark'ının olduğu açık kaynak Türkçe hukuk RAG datasetı.
🇹🇷 Türkçe Özet — Bu dataset, Türkçe hukuk uygulamaları için sıfırdan üretilmiş açık ve denetlenebilir bir RAG corpus'udur. mevzuat.gov.tr üzerinden alınan 25 ana kanunun madde madde temizlenmiş, chunk'lanmış sürümünü (6.350… See the full description on the dataset page: https://huggingface.co/datasets/mtntasci/turkish-legal-rag.turkish-plu-next-event-predictionHomepage: https://github.com/GGLAB-KU/turkish-plu
easy_turkish_math_reasoning
Easy Turkish Math Reasoning
Dataset Summary
The Easy Turkish Math Reasoning dataset is the first phase of a multi-stage curriculum learning pipeline designed to enhance the reasoning abilities of compact language models. This dataset focuses on elementary-level arithmetic and logic problems in Turkish, serving as a warm-up stage for supervised fine-tuning (SFT).
Use Case
Primarily used for:
Bootstrapping reasoning ability in Turkish for compact LLMs.
Phase 1… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/easy_turkish_math_reasoning.medium_turkish_math_reasoning
Dataset Summary
The Medium Turkish Math Reasoning dataset is Phase 2 of a curriculum learning pipeline to teach compact models multi-step reasoning in Turkish. It includes moderately difficult math problems involving multiple reasoning steps, such as two-part arithmetic, comparisons, and logical reasoning.
Use Case
This dataset is ideal for:
Continuing SFT after foundational training with simpler problems.
Bridging the gap between basic arithmetic and complex GSM8K-style… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/medium_turkish_math_reasoning.turkish_llm_finetune_dataset_4_topics
Turkish LLM Finetune Dataset - 4 Topics
This dataset is designed to fine-tune the T3 AI Turkish LLM. It was created by Barathan Aslan, Ömer Faruk Çelik, and Batuhan Kalem for the T3 AI Hackathon. The dataset focuses on four distinct topics: Agriculture, Sustainability, Turkish Education Sytem, and Turkish Law System.
Contributors
Barathan Aslan (https://huggingface.co/barathanasln)
Batuhan Kalem(https://huggingface.co/Pancarsuyu)
Ömer Faruk Çelik… See the full description on the dataset page: https://huggingface.co/datasets/barathanasln/turkish_llm_finetune_dataset_4_topics.collected-turkish-instructions-v0.1This dataset is the result of merging and cleaning data from the following sources:
Turkish Poems Cleaned
Turkish Reading Comprehension Question Answering Dataset
Stanford ALPaCA Cleaned Turkish Translated
Turkish Poems
Turkish Folk Song Lyrics
The data has been merged and processed for quality and consistency to create this dataset.
Genius-Turkish-Dataset
Turkish Song Lyrics from Genius Dataset
Dataset Description
This dataset contains a comprehensive collection of 44,692 Turkish song lyrics, extracted from the larger "Genius Song Lyrics with Language Information" dataset available on Kaggle. The original 9.07 GB dataset was filtered to include only songs identified with the language code 'tr' (Turkish), making it a clean and focused resource for Turkish Natural Language Processing (NLP) tasks.
[TR] Bu veri seti, Kaggle'da… See the full description on the dataset page: https://huggingface.co/datasets/mustafakemal0146/Genius-Turkish-Dataset.atis-ner-turkishThe ATIS (Airline Travel Information System) Dataset includes spoken queries (i.e., utterances) annotated for the task of slot filling in conversational systems.
This dataset, ATISNER, includes airline spoken queries translated from English to Turkish, customized for Named Entity Recognition.
Train and test splits include 4,978 and 890 sentences, respectively.
Translations are provided by the following study.
Şahinuç, F., Yücesoy, V., & Koç, A. (2020). Intent Classification and Slot Filling… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/atis-ner-turkish.finance-reasoning-turkish
Dataset Card for Turkish Advanced Reasoning Dataset (Finance Q&A)
License
This dataset is licensed under the Academic Use Only License. It is intended solely for academic and research purposes. Commercial use is strictly prohibited. For more details, refer to the LICENSE file.
Citation: If you use this dataset in your research, please cite it as follows:
@dataset{turkish_advanced_reasoning_finance_qa,
title = {Turkish Advanced Reasoning Dataset for Finance Q\&A}… See the full description on the dataset page: https://huggingface.co/datasets/emre/finance-reasoning-turkish.turkish-olive-production
Turkish Olive Production Data
Olive and olive oil production in Türkiye at province, region and variety
level. Türkiye is the world's largest producer of table olives, yet its
production figures have not been available as a single machine-readable set
below the national level. This dataset collects them.
Published by zeytin.net ·
Source repository: github.com/yudumnet/zeytinnet
Configurations
Config
Rows
Contents
provinces
45
2024-25 season by province:… See the full description on the dataset page: https://huggingface.co/datasets/yudumnet/turkish-olive-production.Turkish-Product-Review
Turkish Product Review Dataset
This data is orinally from https://www.win.tue.nl/~mpechen/projects/smm/#Datasets
BibTeX Citation
If you use this dataset, please cite following paper:
@inproceedings{Demirtas2013CrosslingualPD,
title={Cross-lingual polarity detection with machine translation},
author={Erkin Demirtas and Mykola Pechenizkiy},
booktitle={wisdom},
year={2013},
url={https://api.semanticscholar.org/CorpusID:3912960}
}
task_categories:
-… See the full description on the dataset page: https://huggingface.co/datasets/asparius/Turkish-Product-Review.turkish-financial-sentiment-256kBu veri seti matriksdata internet sitesinden alınan veriler kullanılarak oluşturulmuştur. Veri setinde Türkiye ekonomisine ilişkin haber metinleri yer almaktadır. Toplamda yaklaşık 256.000 satırdan oluşan veri seti tek parça olarak (train sunulmuştur. Eklenen 'label' sütununda haberin içeriğinin duygu analizini gösteren Negatif, Nötr veya Pozitif değerleri bulunmaktadır. Veri setinin ticari amaçlarla kullanılması tamamen kullanıcıların sorumluluğundadır. Veri seti ile ilgili iletişime geçmek… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/turkish-financial-sentiment-256k.turkish-fake-news-detection
TR-FakeNews: Turkish Fake News Detection on Mainstream Media Dataset
This dataset contains 5325 news title and summaries related to significant events in Türkiye between 2015 and 2023.
Data Fields
title: a string format of the news headline.
description: a string format of the news summary.
status: a classification result 0 (fake) or 1 (real).
updated_log: Information about the data transformation process.
Resources: Indicates the source of the news.
Data Size… See the full description on the dataset page: https://huggingface.co/datasets/isakulaksiz/turkish-fake-news-detection.finance-reasoning-turkish
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti emre (Davut Emre Tasar, Enes Bulut) tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: emre/finance-reasoning-turkish
🔗 Derleyen Platform: VeriPazarı
Türkçe Gelişmiş Akıl Yürütme Veri Seti (Finans Soru-Cevap)
Lisans
Bu veri seti Sadece Akademik… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/finance-reasoning-turkish.
