ruvimx/UkrLM-social
UkrLM Social Corpus A curated corpus of Ukrainian-language text collected from public social media platforms, designed for language model pretraining and fine-tuning. This dataset is part of the UkrLM initiative — an open effort to build foundational NLP resources for the Ukrainian language. Overview Property Value Language Ukrainian (uk) Sources Telegram, Reddit License CC BY 4.0 Format Parquet Task Language Modeling Sources… See the full description on the dataset page: https://huggingface.co/datasets/ruvimx/UkrLM-social.
030
Update README.md
filtering
Update README.md
convert to parquet
add dataset card
add ukrcorpus social dataset
initial commit
