CoolFace
Datasetpublic

ruvimx/UkrLM-social

UkrLM Social Corpus A curated corpus of Ukrainian-language text collected from public social media platforms, designed for language model pretraining and fine-tuning. This dataset is part of the UkrLM initiative — an open effort to build foundational NLP resources for the Ukrainian language. Overview Property Value Language Ukrainian (uk) Sources Telegram, Reddit License CC BY 4.0 Format Parquet Task Language Modeling Sources… See the full description on the dataset page: https://huggingface.co/datasets/ruvimx/UkrLM-social.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes30downloads
7 commits on main
3476fad5mo ago

Update README.md

ruvimx
7f7f7a46mo ago

filtering

ruvimx
59c74de6mo ago

Update README.md

ruvimx
40d9da16mo ago

convert to parquet

ruvimx
bd063476mo ago

add dataset card

ruvimx
ab799e36mo ago

add ukrcorpus social dataset

ruvimx
7288ae46mo ago

initial commit

ruvimx