podarok/kobza-cleaned-ua
kobza-cleaned-ua Cleaned Ukrainian-language subset of Goader/kobza dataset with Russian content filtered out. Dataset Details Dataset Description This dataset is a cleaned and filtered version of the Goader/kobza corpus, removing Russian language content to create a pure Ukrainian language dataset suitable for training language models. The original kobza dataset contains ~60B tokens across 97 million documents. This cleaned version maintains ~59B… See the full description on the dataset page: https://huggingface.co/datasets/podarok/kobza-cleaned-ua.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face