CoolFace
Datasetpublic

podarok/kobza-cleaned-ua

kobza-cleaned-ua Cleaned Ukrainian-language subset of Goader/kobza dataset with Russian content filtered out. Dataset Details Dataset Description This dataset is a cleaned and filtered version of the Goader/kobza corpus, removing Russian language content to create a pure Ukrainian language dataset suitable for training language models. The original kobza dataset contains ~60B tokens across 97 million documents. This cleaned version maintains ~59B… See the full description on the dataset page: https://huggingface.co/datasets/podarok/kobza-cleaned-ua.

sourceHugging Facecc-by-4.0updated 9mo agoView on Hugging Face
0likes22downloads
3 commits on main
8f8c8cc9mo ago

Update dataset card with parent dataset reference and complete documentation

podarok
26b2ae89mo ago

Add cleaned Ukrainian kobza dataset (Russian content filtered)

podarok
6dbc52a9mo ago

initial commit

podarok