datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ukr-emotions-binary
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Binary: specifically this dataset contains binary labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-binary.Ukrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Ukrainian-CulturalHeritage-Books.ipfs_ukraine_laws_ir
Ukraine legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_ukraine_laws (revision 14a17aeb15e04a0d4b9e53c0430d84080805bb5d) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Ukraine prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_ukraine_laws_ir.ukrainian-news-2026
Ukrainian News 2026
Ukrainian-language news articles from 20 national outlets, published between
1 January and 28 August 2026. Extracted body text plus metadata.
Two configs. deduplicated is the default — near-duplicates removed, which
is what you want when mixing this with an already-deduplicated pretraining
corpus. raw is the original release, unchanged.
deduplicated (default)
raw
train-mixin
Documents
419,204
429,427
386,477
Characters
0.97B
1.01B
0.88B
Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Goader/ukrainian-news-2026.ukr-emotions-intensity
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Intensity: specifically this dataset contains intensity labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-intensity.question-answering-ukrainianukr-tg-satire
Ukrainian Telegram satire & troll posts
A corpus of 92k posts from 17 public Telegram channels in the
satire/irony/parody register, harvested July 2026 via the public Telegram
API (history reaching back to 2018 for some channels). The channels are
Ukrainian-audience but the posts are a Ukrainian/Russian mix (the
main truha and sria_news feeds write mostly in Russian, the regional
truexa* branches mostly in Ukrainian), so the corpus carries
language: [ru, uk].
Two flavors are… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/ukr-tg-satire.toxicchat_output-Ukrukr-tg-media
Ukrainian Telegram media posts
A large corpus of Ukrainian news posts harvested from 23 public
Telegram channels (July 2026 snapshot; per-channel history reaching back
to 2016 for some). Channels include national and regional media
(BBC Ukrainian, Zaxidnet, Suspilne News, Nexta, Censor.net, Ukrainska
Pravda, Interfax-Ukraine, Texty, Hromadske, etc.), regional outlets
(Huyovy Kharkiv, Kyiv Real, Oko, Lacheny T, KSPZSU…), and official
accounts (Ministry of Defence of Ukraine, V.… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/ukr-tg-media.recruitment-dataset-candidate-profiles-ukrainian
Djinni Dataset (Ukrainian CVs part)
Overview
The Djinni Recruitment Dataset (Ukrainian CVs part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to candidate CVs, including position titles, candidate information, candidate highlights, job search preferences, job profile types, English… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/recruitment-dataset-candidate-profiles-ukrainian.ukr-emotions-per-annotator
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Per annotator: specifically this dataset contains concatenated… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-per-annotator.yodas-ukrYODAS dataset has both human-annotated and machine-generated transcriptions.
Human-annotated transcriptions are stored in the uk_000.parquet. Transcriptions were artificially shortened which wasn't suitable for LLM pretraining. The shortening is reversed in the stiched version uk_000_stitched.parquet.
Data source
https://huggingface.co/datasets/espnet/yodas
Considerations for Using the Data
Social Impact
This dataset was created to support Ukrainian language… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/yodas-ukr.uk_retail_store_synthetic_dataset
Synthetic Data Generation Demo — UK Retail Dataset
Welcome to this synthetic data generation demo repository by Syncora.ai. This project showcases how to generate synthetic data using real-world tabular structures, demonstrated on a UK retail dataset with columns such as:
Country
CustomerID
UnitPrice
InvoiceDate
Quantity
StockCode
This dataset is designed for dataset for LLM training and AI development, enabling developers to work with privacy-safe, high-quality… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/uk_retail_store_synthetic_dataset.ukrainian-math-openinstructUkrainia_dataset_duplicates
Entity Duplicate Candidates
A dataset of candidate duplicate entity pairs with features and labels, sorted by original.
Columns
original (string): Source entity surface form.
candidate (string): Candidate entity surface form.
levenshtein (float/int): Levenshtein distance (or normalized score if applicable).
normalized (string/bool): Normalized surface form or flag (as in your source).
hashtag (string/bool): Hashtag feature (as in your source).
cossim_score (float):… See the full description on the dataset page: https://huggingface.co/datasets/Jourdain/Ukrainia_dataset_duplicates.dumy-zno-ukrainian-math-history-geo-r1-o1
DUMY («Думи»): Ukrainian Multidomain Reasoning Dataset (Part 1: ZNO/NMT tasks with DeepSeek R1 and OpenAI o1 answers)
DUMY is an open benchmark and dataset designed for training, distillation, and evaluation of language models focused on Ukrainian reasoning tasks.
The word “Dumy” comes from Taras Shevchenko’s famous poem and literally means “thoughts” in Ukrainian:
Думи мої, думи мої,
Лихо мені з вами!
Нащо стали на папері
Сумними рядами?..
Work in progress. Stay tuned.… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/dumy-zno-ukrainian-math-history-geo-r1-o1.echr-ukr-verdict-free
ECtHR Ukraine Verdict-Free Dataset
2,619 verdict-free case-article pairs from the European Court of Human Rights (ECtHR) involving Ukraine as respondent state. Each record contains the factual background, procedural history, parties' arguments, and cited legal standards -- with the Court's legal reasoning, findings, conclusions, and operative provisions surgically removed.
Dataset Description
This dataset is designed for evaluating LLM reliability in legal… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/echr-ukr-verdict-free.ukrainian_ogienko_uk
Ukrainian Bible (Ogienko)
Description
The Ukrainian translation of the Bible by Ivan Ogienko (Митрополит Іларіон), a prominent Ukrainian Orthodox scholar and Metropolitan. First published in 1937, this translation from the original Hebrew and Greek is the classic Ukrainian Bible translation, known for its literary quality and fidelity to the original texts. It includes the Protestant canon (66 books).
Dataset Structure
Each row represents one Bible… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/ukrainian_ogienko_uk.Ukrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/BuzzBlitz360A/Ukrainian-CulturalHeritage-Books.ukrainian_kulish_uk
Ukrainian NT (P. Kulish, 1871)
Description
The Ukrainian New Testament translated by Panteleimon Kulish (1819-1897), a prominent Ukrainian writer, historian, and ethnographer. Published in 1871, this was one of the first complete Ukrainian translations of the New Testament. Kulish translated from the original Greek, using the vernacular Ukrainian of his time. His translation is a landmark of Ukrainian literature and played a significant role in the development of… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/ukrainian_kulish_uk.UKRetailPatterns
UKRetailPatterns
tags: trend analysis, seasonality, forecasting
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'UKRetailPatterns' dataset is designed to assist Machine Learning practitioners in understanding and predicting retail sales trends in the United Kingdom. It includes various factors that could influence retail sales, such as historical sales data, seasonal indicators, and consumer trend analysis. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/UKRetailPatterns.ukri-funding-opportunities
UKRI Funding Opportunities
A structured dataset of 2,102 funding opportunities published by UK Research and Innovation (UKRI) across all nine constituent bodies (AHRC, BBSRC, EPSRC, ESRC, IUK, MRC, NERC, RE, STFC, and cross-council UKRI calls).
This dataset accompanies the paper "Demystifying Funding: Reconstructing a Unified Dataset of the UK Funding Lifecycle" (NSLP 2026) and the GtR Database repository.
Background
UKRI publishes funding opportunities on its website… See the full description on the dataset page: https://huggingface.co/datasets/wrmthorne/ukri-funding-opportunities.Sent_anal_ukr_multi_manualThis dataset is the extension of the existing dataset Sent_anal_ukr_manual (which was made for binary classification). It includes negative (0), positive (1) and neutral (2) sentences, labeled manually. Part of sentences were taken from the book "Francesca. The Queen of Trajectories" by Dorje Batuu (Andriy Vasyliev), a modern Ukrainian author.
sent_anal_ukr_tzpThis is a marked dataset for Ukrainian language. It consists of sentences marked 0, 1 or 2 for negative, neutral or positive mode respectively. The dataset is based on the classic text Shadows of Forgotten Ancestors written by Mykhailo Kotsiubynsky. The markup of the sentences was done automatically based on the lists of positive and negative words from the Sentiment Lexicons for All Major Languages project (Chen & Skiena, ACL 2014). These lists were checked and edited manually by me to… See the full description on the dataset page: https://huggingface.co/datasets/SergiiGurbych/sent_anal_ukr_tzp.sent_anal_ukr_binaryThis dataset for Ukrainian language contains 200 original sentences marked manually with 0 (negative) and 1 (positive).
ukrainian_kulish_puliuy_uk
Ukrainian Bible (Kulish-Puliuy, 1905)
Description
The Kulish-Puliuy Bible (Куліш-Пулюй) is a Ukrainian translation of the Bible prepared by Panteleimon Kulish (1819-1897) and Ivan Puliuy (1845-1918), with the assistance of Ivan Nechuy-Levytsky. Published in 1905, this translation from the original Hebrew and Greek is a landmark of Ukrainian biblical literature. It represents a different translation tradition from the Ogienko version already in the collection. It… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/ukrainian_kulish_puliuy_uk.ukr_to_pt_json_gptSa_Russia_Ukrain_war
