datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rukopys
RUKOPYS: Ukrainian Handwritten Text Recognition Dataset
RUKOPYS (Ukrainian: рукопис — manuscript) is the first large-scale open dataset for Ukrainian handwritten text recognition (HTR). It spans over a century of Ukrainian handwriting — from 1920s archival documents to present-day school homework — and is designed for end-to-end document understanding: region detection, type classification, and text transcription.
Ukrainian is among the largest Slavic languages (45M+ native… See the full description on the dataset page: https://huggingface.co/datasets/UkrainianCatholicUniversity/rukopys.ukraine-forest-change-detection
Satellite Forest Change Detection Dataset
This dataset contains paired multispectral satellite image tiles and binary
change masks for forest change detection. Each valid sample is a triplet:
A: image at the first time point
B: image at the second time point
label: change mask for the same tile and period
The data is organized into train, val, and test splits. Each split has
the same directory structure:
train/
A/
B/
label/
val/
A/
B/
label/
test/
A/
B/
label/… See the full description on the dataset page: https://huggingface.co/datasets/Moriae/ukraine-forest-change-detection.ukrainian-court-decisions
Ukrainian Court Decisions — Judgment Prediction
A dataset of Ukrainian court decisions for case outcome prediction, extracted from the State Court Decisions Registry (ЄДРСР).
Task
Given the facts section (ВСТАНОВИВ) of a court decision, predict the judgment outcome:
Label
Ukrainian
Description
approved
Задоволено
Claim fully satisfied
dismissed
Відмовлено
Claim dismissed
partial
Частково задоволено
Claim partially satisfied
Data… See the full description on the dataset page: https://huggingface.co/datasets/lvanews/ukrainian-court-decisions.UK-Road-DashCamukrainian-court-decisions
Ukrainian Court Decisions — Judgment Prediction
A dataset of Ukrainian court decisions for case outcome prediction, extracted from the State Court Decisions Registry (ЄДРСР).
Task
Given the facts section (ВСТАНОВИВ) of a court decision, predict the judgment outcome:
Label
Ukrainian
Description
approved
Задоволено
Claim fully satisfied
dismissed
Відмовлено
Claim dismissed
partial
Частково задоволено
Claim partially satisfied
Data… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/ukrainian-court-decisions.ukr-emotions-binary
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Binary: specifically this dataset contains binary labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-binary.ukrainian-llm-leaderboard-resultsUK-Road-Bend-Classificationukrainian-newsUkrainian News Dataset
This is a dataset of news articles downloaded from various Ukrainian websites and Telegram channels. The dataset contains approximately ~23M JSON objects (news)EU_acts_in_Ukrainian
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/19753
Description
It was based on: a) the translations of the EU acts in Ukrainian that are available at the official web-portal of the Parliament of Ukraine
https://zakon.rada.gov.ua/laws/main/en/g22), and b) the EU acts that are available in many CEF languages at https://eur-lex.europa.eu.It is a collection of TMX files (X-UK, where X is a CEF language) and includes 3056791 TUs in total.
bg-uk… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/EU_acts_in_Ukrainian.damage_assessment_ukraine
Datasheet
Motivation
For what purpose was the dataset created?
The dataset was created to support research on the impact of the war on Ukrainian infrastructure. It contains satellite and aerial images from before and after disaster events, with the goal of enabling the development and evaluation of models for automated damage assessment.
There is a notable lack of publicly available, labeled datasets representing real-world post-disaster scenarios in Ukraine —… See the full description on the dataset page: https://huggingface.co/datasets/KOlegaBB/damage_assessment_ukraine.Ukrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Ukrainian-CulturalHeritage-Books.details_Radu1999__Mistral-Instruct-Ukrainian-SFT-DPO
Dataset Card for Evaluation run of Radu1999/Mistral-Instruct-Ukrainian-SFT-DPO
Dataset automatically created during the evaluation run of model Radu1999/Mistral-Instruct-Ukrainian-SFT-DPO on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Radu1999__Mistral-Instruct-Ukrainian-SFT-DPO.toronto-tv-ukrainian
Toronto TV Ukrainian Speech Dataset
For educational purposes only.
All rights to the original video and audio content belong to Телебачення Торонто (YouTube channel).
Dataset Summary
A Ukrainian-language speech dataset parsed from the Телебачення Торонто YouTube channel. Each sample consists of a short audio clip and its corresponding Ukrainian subtitle text, intended for use in automatic speech recognition (ASR) research and education.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yuriilaba/toronto-tv-ukrainian.BIRD-UKR
BIRD-UKR: The First Ukrainian Text-to-SQL Benchmark
Overview
BIRD-UKR is the first comprehensive Ukrainian-language benchmark for evaluating NL2SQL (Natural Language to SQL) capabilities of AI models. It is inspired by the English-language BIRD benchmark and fully adapted for Ukrainian — including Ukrainian-language natural language queries, Ukrainian-localized database schemas, and culturally relevant data domains.
Ukrainian remains severely underrepresented in NL2SQL… See the full description on the dataset page: https://huggingface.co/datasets/leev1tan/BIRD-UKR.ipfs_ukraine_laws_ir
Ukraine legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_ukraine_laws (revision 14a17aeb15e04a0d4b9e53c0430d84080805bb5d) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Ukraine prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_ukraine_laws_ir.arc-challenge_ukrtrivia_qa_ukrarc-easy_ukrukrainian-news-2026
Ukrainian News 2026
Ukrainian-language news articles from 20 national outlets, published between
1 January and 28 August 2026. Extracted body text plus metadata.
Two configs. deduplicated is the default — near-duplicates removed, which
is what you want when mixing this with an already-deduplicated pretraining
corpus. raw is the original release, unchanged.
deduplicated (default)
raw
train-mixin
Documents
419,204
429,427
386,477
Characters
0.97B
1.01B
0.88B
Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Goader/ukrainian-news-2026.ipfs_ukraine_laws
Ukraine Laws and Constitution (Verkhovna Rada / zakon.rada.gov.ua)
Research snapshot of official national legislation from Verkhovna Rada / Законодавство України (zakon.rada.gov.ua / data.rada.gov.ua).
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-03
Coverage
snapshot
Source
Verkhovna Rada / Законодавство України (zakon.rada.gov.ua / data.rada.gov.ua)
Collector… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_ukraine_laws.ifeval_ukrukr-toxicity-dataset-seminatural
Ukrainian Toxicity Dataset (Semi-natural)
This is the first of its kind toxicity classification dataset for the Ukrainian language. The datasets was obtained semi-automatically by toxic keywords filtering. For manually collected datasets with crowdsourcing, please, check textdetox/multilingual_toxicity_dataset.
Due to the subjective nature of toxicity, definitions of toxic language will vary. We include items that are commonly referred to as vulgar or profane language. (NLLB paper)… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-toxicity-dataset-seminatural.question-answering-ukrainian-json-answersdetails_SherlockAssistant__Mistral-7B-Instruct-Ukrainian
Dataset Card for Evaluation run of SherlockAssistant/Mistral-7B-Instruct-Ukrainian
Dataset automatically created during the evaluation run of model SherlockAssistant/Mistral-7B-Instruct-Ukrainian on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_SherlockAssistant__Mistral-7B-Instruct-Ukrainian.ukr-emotions-intensity
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Intensity: specifically this dataset contains intensity labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-intensity.ukr-formality-dataset-translated-gyafc
Ukrainian Formality Dataset (translated)
We obtained the first of its kind Ukrainian Formality Classification dataset by trainslating English GYAFC data.
Dataset formation:
English data source: https://aclanthology.org/N18-1012/
Translation into Ukrainian language using model: https://huggingface.co/facebook/nllb-200-distilled-600M
Additionally, the dataset was balanced.
Labels: 0 - informal, 1 - formal.
Load dataset:
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-formality-dataset-translated-gyafc.gsm8k_ukrVOA-ukrukr-dialects-audio-dataset
Ukrainian Dialects Audio Dataset
Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits.
Dataset Description
This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets:
NaUKMA-Audio-Dataset
Ivanna-Stefiuk-Audio-Dataset
Larysa-Irodenko-Audio-Dataset
Hutsulendia-Audio-Dataset
Dido-Yvanchyk-Audio-Dataset-v2
Dataset Structure
train: 27,675 samples
validation: 3… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/ukr-dialects-audio-dataset.
