CoolFace
Datasetpublic

podarok/kobza-cleaned-ua

kobza-cleaned-ua Cleaned Ukrainian-language subset of Goader/kobza dataset with Russian content filtered out. Dataset Details Dataset Description This dataset is a cleaned and filtered version of the Goader/kobza corpus, removing Russian language content to create a pure Ukrainian language dataset suitable for training language models. The original kobza dataset contains ~60B tokens across 97 million documents. This cleaned version maintains ~59B… See the full description on the dataset page: https://huggingface.co/datasets/podarok/kobza-cleaned-ua.

sourceHugging Facecc-by-4.0updated 9mo agoView on Hugging Face
0likes23downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
podarok/kobza-cleaned-ua · CoolFace