CoolFace
Datasetpublic

NumanKaanKaratas/turkish-sentences

Turkish Sentences Turkish Sentences is a clean, duplicate-free Turkish text corpus prepared for NLP and language-model training workflows. The dataset contains Turkish sentences and short lexical entries built around Turkish roots, word forms, homonyms, and morphology-rich vocabulary. Dataset Summary Language: Turkish (tr) Format: Parquet Split: train Rows: 1,978,236 Schema: one column, text Created: 2026-05-31T19:38:26+00:00 Duplicate status: deduplicated Text… See the full description on the dataset page: https://huggingface.co/datasets/NumanKaanKaratas/turkish-sentences.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes50downloads

NumanKaanKaratas/turkish-sentences · main · files are served by the source, never re-hosted here