CoolFace
Datasetpublic

cturan/turkish-synthetic-corpus

Turkish Synthetic Corpus A synthetic Turkish text corpus with 1,871,131 documents, designed for Turkish language model training. About Inspired by HuggingFaceTB/smollm-corpus. Questions and prompts were sourced from the SmolLM Corpus pipeline; a language model then generated localized Turkish responses and documents around them. All credit for the original corpus design and methodology goes to the HuggingFace SmolLM team. The resulting dataset covers a wide range… See the full description on the dataset page: https://huggingface.co/datasets/cturan/turkish-synthetic-corpus.

sourceHugging Faceodc-byupdated 6mo agoView on Hugging Face
0likes76downloads

cturan/turkish-synthetic-corpus · main · files are served by the source, never re-hosted here