Goader/kobza
Kobza On the Path to Make Ukrainian a High-Resource Language [paper] Kobza is the largest publicly available Ukrainian corpus to date, comprising nearly 60 billion tokens across 97 million documents. It is designed to support pretraining and fine-tuning of large language models (LLMs) in Ukrainian, as well as multilingual settings where Ukrainian is underrepresented. 🧾 Dataset Summary Kobza aggregates high-quality Ukrainian text from a wide range of web sources and applies… See the full description on the dataset page: https://huggingface.co/datasets/Goader/kobza.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face