datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ukrainian-stackexchange
Ukrainian StackExchange Dataset
This repository contains a dataset collected from the Ukrainian StackExchange website.
The parsed date is 02/04/2023.
The dataset is in JSON format and includes text data parsed from the website https://ukrainian.stackexchange.com/.
Dataset Description
The Ukrainian StackExchange Dataset is a rich source of text data for tasks related to natural language processing, machine learning, and data mining in the Ukrainian language. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/zeusfsx/ukrainian-stackexchange.ukr-wiki-events
Ukrainian Wikipedia events
A small (1,722-row) Ukrainian dataset built from public-domain /
Wikipedia-sourced text. Two task shapes are mixed in the single train
split (distinguishable via the instruction prompt):
Event extraction — instruction = a passage of Ukrainian Wikipedia
text prefixed by "what important event is this text about:";
output = a short label of the salient event.
Explanation / QA — instruction = a question or term (e.g.
"Опиши явище поліплоїдії"); output =… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/ukr-wiki-events.dumy-zno-ukrainian-math-history-geo-r1-o1
DUMY («Думи»): Ukrainian Multidomain Reasoning Dataset (Part 1: ZNO/NMT tasks with DeepSeek R1 and OpenAI o1 answers)
DUMY is an open benchmark and dataset designed for training, distillation, and evaluation of language models focused on Ukrainian reasoning tasks.
The word “Dumy” comes from Taras Shevchenko’s famous poem and literally means “thoughts” in Ukrainian:
Думи мої, думи мої,
Лихо мені з вами!
Нащо стали на папері
Сумними рядами?..
Work in progress. Stay tuned.… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/dumy-zno-ukrainian-math-history-geo-r1-o1.ukrainian-refugees-financial-advisory
Ukrainian Refugees Financial Advisory Dataset
A dataset of 500 synthetic advisory cases generated by a multi-agent
LLM pipeline that produces and evaluates retirement-oriented financial
guidance for Ukrainian refugee-like profiles in Poland.
Each case covers one full advisory cycle: synthetic profile generation →
draft recommendation + clarifying questions → final structured recommendation
→ automated quality evaluation.
GitHub: uliana0203/ai-agents-refugee-finance… See the full description on the dataset page: https://huggingface.co/datasets/Uliana333/ukrainian-refugees-financial-advisory.
