CoolFace
Datasetpublic

NumanKaanKaratas/turkish-sentences

Turkish Sentences Turkish Sentences is a clean, duplicate-free Turkish text corpus prepared for NLP and language-model training workflows. The dataset contains Turkish sentences and short lexical entries built around Turkish roots, word forms, homonyms, and morphology-rich vocabulary. Dataset Summary Language: Turkish (tr) Format: Parquet Split: train Rows: 1,978,236 Schema: one column, text Created: 2026-05-31T19:38:26+00:00 Duplicate status: deduplicated Text… See the full description on the dataset page: https://huggingface.co/datasets/NumanKaanKaratas/turkish-sentences.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes50downloads
settings

This repository belongs to NumanKaanKaratas on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameturkish-sentences
visibilitypublic
licencemit
gatedno
ownerNumanKaanKaratas
Account settings
NumanKaanKaratas/turkish-sentences · CoolFace