datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rukopys
RUKOPYS: Ukrainian Handwritten Text Recognition Dataset
RUKOPYS (Ukrainian: рукопис — manuscript) is the first large-scale open dataset for Ukrainian handwritten text recognition (HTR). It spans over a century of Ukrainian handwriting — from 1920s archival documents to present-day school homework — and is designed for end-to-end document understanding: region detection, type classification, and text transcription.
Ukrainian is among the largest Slavic languages (45M+ native… See the full description on the dataset page: https://huggingface.co/datasets/UkrainianCatholicUniversity/rukopys.ukrainian-court-decisions
Ukrainian Court Decisions — Judgment Prediction
A dataset of Ukrainian court decisions for case outcome prediction, extracted from the State Court Decisions Registry (ЄДРСР).
Task
Given the facts section (ВСТАНОВИВ) of a court decision, predict the judgment outcome:
Label
Ukrainian
Description
approved
Задоволено
Claim fully satisfied
dismissed
Відмовлено
Claim dismissed
partial
Частково задоволено
Claim partially satisfied
Data… See the full description on the dataset page: https://huggingface.co/datasets/lvanews/ukrainian-court-decisions.ukrainian-court-decisions
Ukrainian Court Decisions — Judgment Prediction
A dataset of Ukrainian court decisions for case outcome prediction, extracted from the State Court Decisions Registry (ЄДРСР).
Task
Given the facts section (ВСТАНОВИВ) of a court decision, predict the judgment outcome:
Label
Ukrainian
Description
approved
Задоволено
Claim fully satisfied
dismissed
Відмовлено
Claim dismissed
partial
Частково задоволено
Claim partially satisfied
Data… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/ukrainian-court-decisions.ukrainian-newsUkrainian News Dataset
This is a dataset of news articles downloaded from various Ukrainian websites and Telegram channels. The dataset contains approximately ~23M JSON objects (news)ukrainian-llm-leaderboard-resultsEU_acts_in_Ukrainian
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/19753
Description
It was based on: a) the translations of the EU acts in Ukrainian that are available at the official web-portal of the Parliament of Ukraine
https://zakon.rada.gov.ua/laws/main/en/g22), and b) the EU acts that are available in many CEF languages at https://eur-lex.europa.eu.It is a collection of TMX files (X-UK, where X is a CEF language) and includes 3056791 TUs in total.
bg-uk… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/EU_acts_in_Ukrainian.Ukrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Ukrainian-CulturalHeritage-Books.details_Radu1999__Mistral-Instruct-Ukrainian-SFT-DPO
Dataset Card for Evaluation run of Radu1999/Mistral-Instruct-Ukrainian-SFT-DPO
Dataset automatically created during the evaluation run of model Radu1999/Mistral-Instruct-Ukrainian-SFT-DPO on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Radu1999__Mistral-Instruct-Ukrainian-SFT-DPO.toronto-tv-ukrainian
Toronto TV Ukrainian Speech Dataset
For educational purposes only.
All rights to the original video and audio content belong to Телебачення Торонто (YouTube channel).
Dataset Summary
A Ukrainian-language speech dataset parsed from the Телебачення Торонто YouTube channel. Each sample consists of a short audio clip and its corresponding Ukrainian subtitle text, intended for use in automatic speech recognition (ASR) research and education.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yuriilaba/toronto-tv-ukrainian.ukrainian-news-2026
Ukrainian News 2026
Ukrainian-language news articles from 20 national outlets, published between
1 January and 28 August 2026. Extracted body text plus metadata.
Two configs. deduplicated is the default — near-duplicates removed, which
is what you want when mixing this with an already-deduplicated pretraining
corpus. raw is the original release, unchanged.
deduplicated (default)
raw
train-mixin
Documents
419,204
429,427
386,477
Characters
0.97B
1.01B
0.88B
Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Goader/ukrainian-news-2026.question-answering-ukrainian-json-answersdetails_SherlockAssistant__Mistral-7B-Instruct-Ukrainian
Dataset Card for Evaluation run of SherlockAssistant/Mistral-7B-Instruct-Ukrainian
Dataset automatically created during the evaluation run of model SherlockAssistant/Mistral-7B-Instruct-Ukrainian on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_SherlockAssistant__Mistral-7B-Instruct-Ukrainian.question-answering-ukrainianukrainian-tts-audiobooks-24khz
Ukrainian Audiobook TTS Dataset (24 kHz)
Description
Ukrainian speech dataset for TTS and ASR tasks.
Source Dataset
https://huggingface.co/datasets/Yehor/audiobooks-xxl
Processing Pipeline
MusicDetection filtering — removed samples with background music/noise
Audio processing (Sidon) — resampled 16 kHz → 24 kHz, converted to mono
Transcription — generated with nvidia/canary-1b-v2
Dataset Structure
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/Mikhailo/ukrainian-tts-audiobooks-24khz.ukrainian-visual-wsd-benchmark
Ukrainain Visual Word Sense Disambiguation Benchmark
Dataset Overview
This dataset is designed for the task of Visual Word Sense Disambiguation, where the goal is to identify, with minimal contextual information, the most appropriate representation of a given ambiguous word from a set of ten images.
Dataset Structure
The dataset is organized into folders, where each folder corresponds to a specific word sense. Each folder contains:
9 images labeled as… See the full description on the dataset page: https://huggingface.co/datasets/yuriilaba/ukrainian-visual-wsd-benchmark.ukrainian-poems
Dataset Card for "ukrainian-poems"
More Information needed
recruitment-dataset-candidate-profiles-ukrainian
Djinni Dataset (Ukrainian CVs part)
Overview
The Djinni Recruitment Dataset (Ukrainian CVs part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to candidate CVs, including position titles, candidate information, candidate highlights, job search preferences, job profile types, English… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/recruitment-dataset-candidate-profiles-ukrainian.ukrainian-stackexchange
Ukrainian StackExchange Dataset
This repository contains a dataset collected from the Ukrainian StackExchange website.
The parsed date is 02/04/2023.
The dataset is in JSON format and includes text data parsed from the website https://ukrainian.stackexchange.com/.
Dataset Description
The Ukrainian StackExchange Dataset is a rich source of text data for tasks related to natural language processing, machine learning, and data mining in the Ukrainian language. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/zeusfsx/ukrainian-stackexchange.autotrain-data-ukrainian-telegram-sentiment-analysis
AutoTrain Dataset for project: ukrainian-telegram-sentiment-analysis
Dataset Description
This dataset has been automatically processed by AutoTrain for project ukrainian-telegram-sentiment-analysis.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": "\u0421\u043e\u0432\u043e\u043a",
"target": 1
},
{
"text":… See the full description on the dataset page: https://huggingface.co/datasets/dmytrobaida/autotrain-data-ukrainian-telegram-sentiment-analysis.recruitment-dataset-job-descriptions-ukrainian
Djinni Dataset (Ukrainian Job Descriptions part)
Overview
The Djinni Recruitment Dataset (Ukrainian Job Descriptions part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to job descriptions, including position titles, job descriptions, company names, experience requirements, keywords… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/recruitment-dataset-job-descriptions-ukrainian.gemma3_ukrainian-ner-datasetukrainian-dialogs-constructiveness
Ukrainian Online Discourse: A Two-Axis Constructiveness Dataset
Expert-annotated Ukrainian-language dialogues from Telegram and the Ukrayinska
Pravda Forum, scored on six ordinal dimensions of dialogic constructiveness
organized into two orthogonal axes — Relational Conduct (RC) and
Substantive Contribution (SC).
Note. This dataset accompanies an exploratory study. Findings should be
treated as preliminary, given the single-language setting and the modest
size of the… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/ukrainian-dialogs-constructiveness.ULP-Ukrainian-Language-Proficiency
Ukrainian Language Proficiency (ULP) Benchmark
Dataset Description
The Ukrainian Language Proficiency (ULP) benchmark is an expert-curated dataset designed to evaluate Ukrainian language proficiency in Large Language Models (LLMs), with a focus on grammar and orthography as core components of language competence.
Dataset Summary
This gold-standard dataset contains 347 multiple-choice questions prepared by professional linguists.
The benchmark… See the full description on the dataset page: https://huggingface.co/datasets/SGaleshchuk/ULP-Ukrainian-Language-Proficiency.WizardLM-ukrainian
WizardLM Translated to Ukrainian 🇺🇦
Dataset Description
A Ukrainian language dataset comprising 140,000+ records translated from the WizardLM dataset.
This dataset is suitable for various natural language processing tasks.
This is not merged with original ShareGPT threads.
Data translated via using Google Gemini Pro API.
Слава Україні!
Disclaimer
Prepare data before your usage. There are some errors in texts, so be carefull.
How to Use
This… See the full description on the dataset page: https://huggingface.co/datasets/cidtd-mod-ua/WizardLM-ukrainian.details_Radu1999__Mistral-Instruct-Ukrainian-SFT
Dataset Card for Evaluation run of Radu1999/Mistral-Instruct-Ukrainian-SFT
Dataset automatically created during the evaluation run of model Radu1999/Mistral-Instruct-Ukrainian-SFT on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Radu1999__Mistral-Instruct-Ukrainian-SFT.ukrainian-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction - Ukraine
The Synthetic Ukraine Passports Dataset compiles more than 1,000 AI-generated passport images created for training OCR and computer vision models on identity documents. Each record is fully synthetic, so the… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/ukrainian-passports.sovereign-ukrainian-sft-core
ChronoMatrix: Суверенний українськомовний SFT-корпус експертного рівня
Ручна робота. Академічний фундамент. Промисловий формат.
600+ валідних JSONL-рядків для SFT / fine-tuning українськомовних LLM. Створено вручну, без машинного перекладу й ідеологічних штампів, у форматі, готовому до прямого завантаження в тренувальний конвеєр (Llama, Mistral, Gemma, Qwen тощо).
📋 Один зразок — перед тим, як читати опис
{"text": "<user>Змоделюй сцену: липень 971 року, береги… See the full description on the dataset page: https://huggingface.co/datasets/chronomatrix/sovereign-ukrainian-sft-core.alpaca-ukrainian-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-ukrainian-cleaned.alpaca_ukrainian_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_ukrainian_taco.ukrainian-treebank-lmUkrainian part of the Universal Dependencies, specifically preprocessed for the language modeling task. The data can be split into documents, paragraphs or sentences. Manual selection of the data done by the authors of the dataset makes it suitable for the perplexity evaluation.
Authors of the dataset: Institute for Ukrainian, NGO, org@mova.institute
GitHub: https://github.com/UniversalDependencies/UD_Ukrainian-IU
