CoolFace
Datasetpublic

ytu-ce-cosmos/Cosmos-Turkish-Corpus-v1.0

This is the Turkish pretraining corpus of the Cosmos AI Research Group. It contains ~15B tokens and demonstrates competitive performance across various Turkish benchmarks when used in continual pretraining setups. Cosmos-Turkish-Corpus is collected from a wide range of Turkish websites, including forums, news sources, Wikipedia, and more. URL-based deduplication has been applied; however, additional content-level deduplication and filtering may be required before use.

sourceHugging Facecc-by-4.0updated 10mo agoView on Hugging Face
26likes582downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
ytu-ce-cosmos/Cosmos-Turkish-Corpus-v1.0 · CoolFace