datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ULP-Ukrainian-Language-Proficiency
Ukrainian Language Proficiency (ULP) Benchmark
Dataset Description
The Ukrainian Language Proficiency (ULP) benchmark is an expert-curated dataset designed to evaluate Ukrainian language proficiency in Large Language Models (LLMs), with a focus on grammar and orthography as core components of language competence.
Dataset Summary
This gold-standard dataset contains 347 multiple-choice questions prepared by professional linguists.
The benchmark… See the full description on the dataset page: https://huggingface.co/datasets/SGaleshchuk/ULP-Ukrainian-Language-Proficiency.clean_ukrainian-news
Ukrainian News Dataset
This is a dataset of news articles downloaded from various Ukrainian websites and Telegram channels.
The dataset contains 22 567 099 JSON objects (news), total size ~67GB each with the following fields:
title: The title of the news article
text: The text of the news article, which may contain HTML tags(e.g., paragraphs, links, images, etc.)
url: The URL of the news article
datetime: The time of publication or when the article was parsed and added to… See the full description on the dataset page: https://huggingface.co/datasets/itsSHAS/clean_ukrainian-news.ukrainian-safety-dataset
Ukrainian LLM Safety Dataset
Dataset Summary
A national security-focused safety classification dataset for Ukrainian LLM infrastructure, developed in partnership with the Ministry of Digital Transformation of Ukraine. Contains adversarial prompts across a 9-class threat taxonomy with synthetic rationales and contrastive examples for knowledge distillation.
Files
File
Samples
Purpose
train.jsonl
59,377
Training corpus with contrastive… See the full description on the dataset page: https://huggingface.co/datasets/mark-matviiv/ukrainian-safety-dataset.ukrainian_wsd_benchmark
Ukrainian WSD Benchmark
This dataset links Ukrainian homonym senses to naturally occurring sentences from the General Regionally Annotated Corpus of Ukrainian (GRAC). It contains 1,386 lemmas, 3,206 senses, and 13,310 examples. Each retained lemma has at least two senses with at least one example each.
Construction
The homonyms and sense definitions were selected by professional linguists. The pipeline normalized and deduplicated the inventory, retrieved up to 1… See the full description on the dataset page: https://huggingface.co/datasets/yuriilaba/ukrainian_wsd_benchmark.ukrainian-refugees-financial-advisory
Ukrainian Refugees Financial Advisory Dataset
A dataset of 500 synthetic advisory cases generated by a multi-agent
LLM pipeline that produces and evaluates retirement-oriented financial
guidance for Ukrainian refugee-like profiles in Poland.
Each case covers one full advisory cycle: synthetic profile generation →
draft recommendation + clarifying questions → final structured recommendation
→ automated quality evaluation.
GitHub: uliana0203/ai-agents-refugee-finance… See the full description on the dataset page: https://huggingface.co/datasets/Uliana333/ukrainian-refugees-financial-advisory.
