datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ukrainian-newsUkrainian News Dataset
This is a dataset of news articles downloaded from various Ukrainian websites and Telegram channels. The dataset contains approximately ~23M JSON objects (news)Ukrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Ukrainian-CulturalHeritage-Books.ukrainian-news-2026
Ukrainian News 2026
Ukrainian-language news articles from 20 national outlets, published between
1 January and 28 August 2026. Extracted body text plus metadata.
Two configs. deduplicated is the default — near-duplicates removed, which
is what you want when mixing this with an already-deduplicated pretraining
corpus. raw is the original release, unchanged.
deduplicated (default)
raw
train-mixin
Documents
419,204
429,427
386,477
Characters
0.97B
1.01B
0.88B
Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Goader/ukrainian-news-2026.WizardLM-ukrainian
WizardLM Translated to Ukrainian 🇺🇦
Dataset Description
A Ukrainian language dataset comprising 140,000+ records translated from the WizardLM dataset.
This dataset is suitable for various natural language processing tasks.
This is not merged with original ShareGPT threads.
Data translated via using Google Gemini Pro API.
Слава Україні!
Disclaimer
Prepare data before your usage. There are some errors in texts, so be carefull.
How to Use
This… See the full description on the dataset page: https://huggingface.co/datasets/cidtd-mod-ua/WizardLM-ukrainian.ukrainian-treebank-lmUkrainian part of the Universal Dependencies, specifically preprocessed for the language modeling task. The data can be split into documents, paragraphs or sentences. Manual selection of the data done by the authors of the dataset makes it suitable for the perplexity evaluation.
Authors of the dataset: Institute for Ukrainian, NGO, org@mova.institute
GitHub: https://github.com/UniversalDependencies/UD_Ukrainian-IUdumy-zno-ukrainian-math-history-geo-r1-o1
DUMY («Думи»): Ukrainian Multidomain Reasoning Dataset (Part 1: ZNO/NMT tasks with DeepSeek R1 and OpenAI o1 answers)
DUMY is an open benchmark and dataset designed for training, distillation, and evaluation of language models focused on Ukrainian reasoning tasks.
The word “Dumy” comes from Taras Shevchenko’s famous poem and literally means “thoughts” in Ukrainian:
Думи мої, думи мої,
Лихо мені з вами!
Нащо стали на папері
Сумними рядами?..
Work in progress. Stay tuned.… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/dumy-zno-ukrainian-math-history-geo-r1-o1.clean_ukrainian-news
Ukrainian News Dataset
This is a dataset of news articles downloaded from various Ukrainian websites and Telegram channels.
The dataset contains 22 567 099 JSON objects (news), total size ~67GB each with the following fields:
title: The title of the news article
text: The text of the news article, which may contain HTML tags(e.g., paragraphs, links, images, etc.)
url: The URL of the news article
datetime: The time of publication or when the article was parsed and added to… See the full description on the dataset page: https://huggingface.co/datasets/itsSHAS/clean_ukrainian-news.Ukrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/BuzzBlitz360A/Ukrainian-CulturalHeritage-Books.ukrainian-refugees-financial-advisory
Ukrainian Refugees Financial Advisory Dataset
A dataset of 500 synthetic advisory cases generated by a multi-agent
LLM pipeline that produces and evaluates retirement-oriented financial
guidance for Ukrainian refugee-like profiles in Poland.
Each case covers one full advisory cycle: synthetic profile generation →
draft recommendation + clarifying questions → final structured recommendation
→ automated quality evaluation.
GitHub: uliana0203/ai-agents-refugee-finance… See the full description on the dataset page: https://huggingface.co/datasets/Uliana333/ukrainian-refugees-financial-advisory.ukrainian-newsUkrainian News Dataset
This is a dataset of news articles downloaded from various Ukrainian websites and Telegram channels. The dataset contains approximately ~23M JSON objects (news)
