CoolFace
Datasetpublic

Goader/kobza

Kobza On the Path to Make Ukrainian a High-Resource Language [paper] Kobza is the largest publicly available Ukrainian corpus to date, comprising nearly 60 billion tokens across 97 million documents. It is designed to support pretraining and fine-tuning of large language models (LLMs) in Ukrainian, as well as multilingual settings where Ukrainian is underrepresented. 🧾 Dataset Summary Kobza aggregates high-quality Ukrainian text from a wide range of web sources and applies… See the full description on the dataset page: https://huggingface.co/datasets/Goader/kobza.

sourceHugging Facecc0-1.0updated 1y agoView on Hugging Face
14likes278downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Goader/kobza · CoolFace