datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-2-turkish-categorized
What is this
THis is the categorized version of the Turkish subset of the fineweb-2 dataset.
It is an ongoing effort, and the details will be added soon with the rest of the dataset.
turkishfineweb2-cleaned
TurkishFineweb2-Cleaned
A Turkish web corpus derived from the Turkish (tur_Latn) subset of
FineWeb-2, augmented with an additional quality-classification layer
and a near-duplicate removal pass.
📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum
Source
FineWeb-2 is a
large-scale, multilingual web corpus built from Common Crawl. This dataset
covers the Turkish (tur_Latn) portion of FineWeb-2, spanning the… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkishfineweb2-cleaned.Turkish_corpus
Turkish Corpus 🇹🇷
Turkish Corpus is a large-scale cleaned Turkish text dataset created by collecting public Turkish corpora and extracting Turkish-language portions from multilingual datasets. The dataset is designed for Turkish Natural Language Processing research, language model pretraining, tokenizer training, embedding models, retrieval systems, and general Turkish language understanding tasks.
The main purpose of this dataset is to provide a practical, scalable, and… See the full description on the dataset page: https://huggingface.co/datasets/Ethosoft/Turkish_corpus.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/mrfg/turkish-court-decisions.turkish_embedding_model_training_datamogan-turkish-web
Mogan Turkish Web
A large-scale Turkish web corpus derived from monthly Common Crawl snapshots
covering the period from January 2025 to June 2026. The corpus was
constructed by extracting Turkish-language content from raw Common Crawl
WARC/WET dumps, followed by language filtering, PII masking, and
near-duplicate removal.
📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum
Dataset Summary
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/moganai/mogan-turkish-web.turkish-corpus-100b
Turkish Corpus 100B (TC-100B)
Dataset Summary
The Turkish Corpus 100B (TC-100B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Turkish. Comprising approximately 105 Billion tokens (measured with Qwen/Llama3 tokenizer), it represents one of the largest open resources for Turkish LLM pretraining.
The dataset is engineered for a two-stage training pipeline:
Pretrain Subset (~103B Tokens): A diverse mix of… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/turkish-corpus-100b.turkish-tts-combined-raw
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından düzenlenmiştir. Orijinal veri seti afkfatih tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: afkfatih/turkish-tts-combined-raw
🔗 Derleyen Platform: VeriPazarı
Türkçe TTS Birleşik Veri Seti (Turkish TTS Combined)
7 farklı açık kaynak Türkçe TTS (Metinden Sese) veri setinin birleşimidir.
~81.500 örnek |… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/turkish-tts-combined-raw.fineweb2-hq-turkishCosmos-Turkish-Corpus-v1.0This is the Turkish pretraining corpus of the Cosmos AI Research Group.
It contains ~15B tokens and demonstrates competitive performance across various Turkish benchmarks when used in continual pretraining setups.
Cosmos-Turkish-Corpus is collected from a wide range of Turkish websites, including forums, news sources, Wikipedia, and more.
URL-based deduplication has been applied; however, additional content-level deduplication and filtering may be required before use.
turkish-daily-dialogues-5k
Turkish Daily Dialogues 5K
Exactly 5,000 synthetic, multi-turn Turkish conversations covering ordinary daily-life situations. The corpus is designed as a small, auditable baseline for dialogue modelling, instruction-format experiments, augmentation research, and Turkish-language evaluation—not as a substitute for conversations written by real people.
Provenance in one sentence: the Turkish source scenario library was drafted with AI assistance specifically for this project… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/turkish-daily-dialogues-5k.image-captioning-turkish
Türkçe Image Captioning Veri Seti
Bu veri seti BLIP3o modelinin pretrain eğitiminde kullanılan BLIP3o-Pretrain-Long-Caption ve BLIP3o-Pretrain-Short-Caption veri setlerinin Türkçeye çevirilmiş bir alt parçasıdır. Orijinal veri setinin oluşturulması ile ilgili detaylı bilgiye BLIP-3o makalesi üzerinden ulaşabilirsiniz.
Veri seti Image-to-Text modellerinin eğitilmesinde veya ince ayar sürecinde kullanılabilir. Veri seti, orijinal veri setinin lisansı olan Apache 2.0 altında… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/image-captioning-turkish.turkish_corpusGoogle-Translated_Turkish_GPQA_Datasetcleaned_turkish_embedding_model_training_data_colabTurkish_TTS_Datacleaned_turkish_embedding_model_training_data_colab
Citation
If you use this dataset in your research, please cite the following paper:
@inproceedings{baysan-gungor-2025-tr,
title = "{TR}-{MTEB}: A Comprehensive Benchmark and Embedding Model Suite for {T}urkish Sentence Representations",
author = "Baysan, Mehmet Selman and
Gungor, Tunga",
booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2025",
month = nov,
year = "2025",
address = "Suzhou, China",
publisher =… See the full description on the dataset page: https://huggingface.co/datasets/trmteb/cleaned_turkish_embedding_model_training_data_colab.turkish_corpus_tokenized
Dataset Card for "turkish_corpus_tokenized"
More Information needed
turkish-tts-combined-raw
Türkçe TTS Birleşik Veri Seti
7 farklı açık kaynak Türkçe TTS veri setinin birleşimi. ~81,500 örnek | 24kHz | SNAC uyumlu
Kaynaklar
Veri Seti
Örnek
Kaynak
Mazlum Kiper
9,643
omersaidd/tts_mazlum_kiper_tur
Ahmet Deniz
11,289
omersaidd/tts_ahmet_deniz_tur
Nisan Kumru
8,042
omersaidd/tts_nisan_kumru_tur
Derya TTS v2
42
afkfatih/derya-tts-v2
Derya Karma v3
255
afkfatih/derya-tts-karma-v3
Khan Academy
25,741
ysdede/khanacademy-turkish
Common… See the full description on the dataset page: https://huggingface.co/datasets/projectkaira/turkish-tts-combined-raw.turkish_parliamentary_data
Grand National Assembly Corpus of Türkiye (GNACT)
A comprehensive collection of Turkish parliamentary transcripts spanning over 100 years (1920–present), from 10 legislative bodies. Includes both Ottoman Turkish (1920–1928) and Modern Turkish (1928–present) texts.
Loading the dataset
from datasets import load_dataset
# Strategy 1: full session documents, all bodies (default)
ds = load_dataset("boun-tabilab/turkish_parliamentary_data", "full_sessions", split="train")
#… See the full description on the dataset page: https://huggingface.co/datasets/boun-tabilab/turkish_parliamentary_data.turkish-tts-combined-raw
Türkçe TTS Birleşik Veri Seti
7 farklı açık kaynak Türkçe TTS veri setinin birleşimi. ~81,500 örnek | 24kHz | SNAC uyumlu
Kaynaklar
Veri Seti
Örnek
Kaynak
Mazlum Kiper
9,643
omersaidd/tts_mazlum_kiper_tur
Ahmet Deniz
11,289
omersaidd/tts_ahmet_deniz_tur
Nisan Kumru
8,042
omersaidd/tts_nisan_kumru_tur
Derya TTS v2
42
afkfatih/derya-tts-v2
Derya Karma v3
255
afkfatih/derya-tts-karma-v3
Khan Academy
25,741
ysdede/khanacademy-turkish
Common Voice 17
26… See the full description on the dataset page: https://huggingface.co/datasets/afkfatih/turkish-tts-combined-raw.turkish_clip_dataset_with_text_embeddingsThis dataset cleaned and dowloaded version of following dataset: https://huggingface.co/datasets/visheratin/laion-coco-nllb
The main purpose was to extract Turkish captions and download images.
You can use this dataset to fine-tune or create a clip model.
Since there English and Turkish captions you can also use those to create language model?
turkish-tts-audiobooks
Turkish TTS Audiobooks
Turkish read-speech corpus for text-to-speech training, built from Turkish
audiobook and spoken-article recordings by an automatic pipeline: VAD
segmentation → technical QC → acoustic event tagging → DNSMOS → speaker
embedding/consistency → double-pass Whisper ASR → text policy → leakage-free
splitting. Audio is 16 kHz mono lossless FLAC embedded in the Parquet shards.
The pipeline that produced it — every stage, every threshold, the export and
audit… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/turkish-tts-audiobooks.toksuite_turkish
Dataset Card for Tokenization Robustness
TokSuite Benchmark (Turkish Collection)
Dataset Description
This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Turkish language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness.
Curated by: R3 Research Team… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_turkish.turkish-nli-constructed-1.5m
Turkish NLI Constructed 1.5M v2
Dengeli entailment, contradiction ve neutral sınıflı Türkçe NLI çiftleri.
Doğrulanmış boyut
Train: 1,470,000
Validation: 15,000
Test: 15,000
Toplam: 1,500,000
Ana görev sütunları: id, premise, hypothesis, label
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-nli-constructed-1.5m.turkish-court-decisions
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/Alptekinege/turkish-court-decisions.turkishvoicedataset
Dataset Card for "turkishneuralvoice"
Dataset Overview
Dataset Name: Turkish Neural Voice
Description: This dataset contains Turkish audio samples generated using Microsoft Text to Speech services. The dataset includes audio files and their corresponding transcriptions.
Dataset Structure
Configs:
default
Data Files:
Split: train
Path: data/train-*
Dataset Info:
Features:
audio: Audio file
transcription: Corresponding text transcription
Splits:
train… See the full description on the dataset page: https://huggingface.co/datasets/erenfazlioglu/turkishvoicedataset.image-vqa-turkish
Türkçe Image VQA Veri Seti
Bu veri seti Türkçe görsel soru-cevap (VQA) çiftleri içermektedir.
Kullanım
from datasets import load_dataset
ds = load_dataset("ituperceptron/turkish-image-vqa", split="vqa")
Veri Yapısı
image: Görsel (PIL Image)
image_id: Görselin benzersiz ID'si
vqa: VQA soru-cevap çiftleri (JSON formatında)
turkish_product_reviews
Dataset Card for Turkish Product Reviews
Dataset Summary
This Turkish Product Reviews Dataset contains 235.165 product reviews collected online. There are 220.284 positive, 14881 negative reviews.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The dataset is based on Turkish.
Dataset Structure
Data Instances
Example 1:
sentence: beklentimin altında bir ürün kaliteli değil
sentiment: 0 (negative)
Example 2:… See the full description on the dataset page: https://huggingface.co/datasets/fthbrmnby/turkish_product_reviews.turkish_male
