podarok/kobza-cleaned-ua
kobza-cleaned-ua Cleaned Ukrainian-language subset of Goader/kobza dataset with Russian content filtered out. Dataset Details Dataset Description This dataset is a cleaned and filtered version of the Goader/kobza corpus, removing Russian language content to create a pure Ukrainian language dataset suitable for training language models. The original kobza dataset contains ~60B tokens across 97 million documents. This cleaned version maintains ~59B… See the full description on the dataset page: https://huggingface.co/datasets/podarok/kobza-cleaned-ua.
Update dataset card with parent dataset reference and complete documentation
Add cleaned Ukrainian kobza dataset (Russian content filtered)
initial commit
