CoolFace
Datasetpublic

ytu-ce-cosmos/Cosmos-Turkish-Corpus-v1.0

This is the Turkish pretraining corpus of the Cosmos AI Research Group. It contains ~15B tokens and demonstrates competitive performance across various Turkish benchmarks when used in continual pretraining setups. Cosmos-Turkish-Corpus is collected from a wide range of Turkish websites, including forums, news sources, Wikipedia, and more. URL-based deduplication has been applied; however, additional content-level deduplication and filtering may be required before use.

sourceHugging Facecc-by-4.0updated 10mo agoView on Hugging Face
26likes582downloads
7 commits on main
8c9a64010mo ago

Update README.md

Yavuz Selim Aygan
416dfb810mo ago

Update README.md

Yavuz Selim Aygan
1b96b6310mo ago

Update README.md

Yavuz Selim Aygan
e5d85ce10mo ago

Upload dataset (part 00002-of-00003)

Yavuz Selim Aygan
756b43310mo ago

Upload dataset (part 00001-of-00003)

Yavuz Selim Aygan
3be646410mo ago

Upload dataset (part 00000-of-00003)

Yavuz Selim Aygan
9047d5e10mo ago

initial commit

Yavuz Selim Aygan