datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cyber-security
Cybersecurity Instruction-Tuning Dataset
A large, cleaned, multi-domain cybersecurity chat dataset for LLM finetuning,
built from 198 distinct sources spanning offensive security, blue-team
operations, vulnerability intelligence, cloud/AWS security, malware analysis,
digital forensics, and more. Every record is normalized to the standard
messages chat format and deduplicated at both file and record level.
⚠️ Research use only. This dataset is provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/cyber-security.service-uahj
传奇私服分布式路由与自动化接口索引库 - Batch 001
本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:传奇私服)。
📂 区域节点集群子目录 (Spider Pool Indexes)
👉 新开传奇私服 - 传奇私服网 - 传奇私服发布 —— 承载源站 🌐 api.agergames.com
👉 传奇私服推荐 - 新开传奇私服 - 传奇私服 —— 承载源站 🌐 api.amoygame.com
👉 今日传奇私服 - 传奇私服推荐 - 传奇私服网 —— 承载源站 🌐 api.gamecctv.com
👉 传奇私服999 - 传奇私服发布网 - 传奇私服 —— 承载源站 🌐 api.gamerling.com
👉 传奇私服发布 - 传奇私服999 - 新开传奇私服 —— 承载源站 🌐 api.games-vr.com
👉 传奇私服 - 今日传奇私服 - 传奇私服999 —— 承载源站 🌐 api.j5games.com
👉 今日传奇私服 - 传奇私服发布 -… See the full description on the dataset page: https://huggingface.co/datasets/makesi1/service-uahj.uzbek-instruct-llmuzbek-instruct-llm is a corpus of more than 15,000 records. It's made for instruct fine-tuning large language models for Uzbek language. It's mostly translated from other instruct datasets with some extra data added
ua-llm-router-eval
UA Specialist Router — evaluation progress
Score tables, routing stats, and run metadata from the diploma project
MariaOnyshchuk/ua-llm-router:
a rules-based router over open Ukrainian specialists (Mamay-4B, Lapa-12B, Aya Expanse 8B, Qwen-Coder).
This dataset is the progress log of pinned JSON summaries, not a dump of every generation.
What is included
Path
Contents
progress_ledger.csv
Flattened metric rows across weeks (best table for browsing)… See the full description on the dataset page: https://huggingface.co/datasets/MariaOnyshchuk/ua-llm-router-eval.tiny-ua-bench-responses
Tiny-UA-Bench Responses
This dataset contains the response matrix for Tiny-UA-Bench.
The matrix contains 919,160 model and item records.
The matrix covers 20 models and 45,958 items.
The evaluation excludes FLORES and LongFLORES.
Use
Use this dataset to reproduce the benchmark compression analysis.
Do not use a held-out model response to fit a selector or predictor.
Use the reference and held-out split definitions from the code repository.
Load the data with the… See the full description on the dataset page: https://huggingface.co/datasets/robinhad/tiny-ua-bench-responses.UCDS
uCoder Dataset
A high-quality, deduplicated dataset for training coding and mathematics language models.
Dataset Statistics
Total Samples: 420,686
Format: ChatML (messages array)
Languages: Python, JavaScript, C++, Java, and more
Sources
This dataset merges and cleans data from:
Source
Samples
Description
ByteDance-Seed/Code-Contests-Plus
10,293
Competitive programming
open-r1/codeforces
47,558
Codeforces problems
sahil2801/CodeAlpaca-20k
15… See the full description on the dataset page: https://huggingface.co/datasets/uaytug/UCDS.fumea-dataset
FUMEA Dataset
FUMEA-Dataset is a merged, curated, and deduplicated corpus designed for Supervised Fine-Tuning (SFT) of large language models. It unifies two specialized domains — tool-use / function-calling and financial analysis — into a single, training-ready resource. All samples are pre-formatted with the Qwen3 chat template (<|im_start|> / <|im_end|>) and require no additional preprocessing.
This dataset is the primary training resource behind the FUMEA-F model family, which… See the full description on the dataset page: https://huggingface.co/datasets/uaytug/fumea-dataset.ua-tg-misc
UA/TG misc — 15 small channel stubs
A bundle of the 15 small Telegram channel exports (10 of them kept non-empty text rows; the rest exported only media-only posts, stripped here) — stubs and partially
scraped channels that did not reach the full 10k-message scrape that the
main corpora (hausmer/ukr-tg-media,
hausmer/ukr-tg-satire)
got. Most are 1–30 text posts (some channels were only reachable briefly,
others export empty media rows).
This corpus contains 67 text posts across… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/ua-tg-misc.Synthetic_UAV_Scenario_LLM_MultiTurn_GPS_Navigation
Dataset Card for CJJones/Synthetic_UAV_Scenario_LLM_MultiTurn_GPS_Navigation
The full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com
Want more? 🚀 Get the AI Startup Bundle from Gumroad.
Dataset Summary
This dataset contains synthetic multi-turn UAV (Unmanned Aerial Vehicle) flight scenarios with realistic GPS navigation challenges, flight mode transitions, and system diagnostics. The scenarios simulate various flight conditions… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Synthetic_UAV_Scenario_LLM_MultiTurn_GPS_Navigation.WizardLM-ukrainian
WizardLM Translated to Ukrainian 🇺🇦
Dataset Description
A Ukrainian language dataset comprising 140,000+ records translated from the WizardLM dataset.
This dataset is suitable for various natural language processing tasks.
This is not merged with original ShareGPT threads.
Data translated via using Google Gemini Pro API.
Слава Україні!
Disclaimer
Prepare data before your usage. There are some errors in texts, so be carefull.
How to Use
This… See the full description on the dataset page: https://huggingface.co/datasets/cidtd-mod-ua/WizardLM-ukrainian.CVEs
CVEs — a full-coverage CVE chat dataset
1,625,017 chat conversations covering all 361,190 usable CVEs (1999–2026), built for fine-tuning cybersecurity assistants. Every known CVE in the official CVE List with severity enrichment from NVD (via the fkie-cad community feeds), rendered as English user/assistant conversations with varied phrasings, honest handling of missing data, and a per-CVE 99/1 train/validation split with zero leakage.
The schema matches oi-uae/cyber-security… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/CVEs.ua-codeforces-cots-open-r1-for-training
Version of anon-researcher-ua/ua-codeforces-cots-open-r1 prepared for model training
UA4RAG-it
UA4RAG
📘 Dataset Summary
UA4RAG (UnAnswerable for RAG) is a collection of datasets designed to train and evaluate language models on generating and recognizing unanswerable factual questions and appropriate non-answers given a reference text.
In retrieval-augmented generation (RAG) systems, retrieved contexts are often tangential to user queries. This dataset addresses the critical challenge of training models to recognize when sufficient evidence is absent and to… See the full description on the dataset page: https://huggingface.co/datasets/lopozz/UA4RAG-it.uae-laws-irac
UAE Laws Q&A Dataset (IRAC Format)
A high-quality dataset of 9,477 question-answer pairs about UAE laws, formatted in IRAC (Issue, Rule, Application, Conclusion) legal reasoning structure.
Dataset Creation
Source Documents
The dataset was built from a comprehensive collection of UAE legal documents, including:
Federal Decrees and Laws
Cabinet Resolutions
Ministerial Decisions
Civil and Commercial Codes
Labor Law
Traffic Law
And more
Creation Process… See the full description on the dataset page: https://huggingface.co/datasets/SalahALHaismawi/uae-laws-irac.ua-legal-citation-grounded-sft
UA Legal Citation-Grounded SFT
A supervised fine-tuning set of citation-grounded legal question-answering examples in
Ukrainian. Every assistant answer attributes each factual claim to a specific source with a
[doc:ID] marker that refers to a real court decision passage placed in the prompt. The set
is built to train and study retrieval-grounded generation where faithfulness of citations,
not just answer quality, is the target.
How it was built
Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/ua-legal-citation-grounded-sft.structured-uae-laws
Dataset Card for structured-uae-laws
This dataset is a collection of question & answers about the laws and regulations in the United Arab Emirates.
It covers different areas of law like:
economy and business
family and community
finance and banking
industry and technical standardisation
justice and juiciary, labour
residency and leberal professions
security and safety
tax
Dataset Sources
Repository
Base Dataset
United Arab Emirates Legislations… See the full description on the dataset page: https://huggingface.co/datasets/obadabaq/structured-uae-laws.ua-codeforces-cots-open-r1
Dataset Summary
ua-codeforces-cots-open-r1 is a Ukrainian-focused derivative of open-r1/codeforces-cots that:
includes 1550 Python solutions from original dataset generated by DeepSeek-R1;
adds Ukrainian translations of Codeforces task statements, I/O formats, notes, and editorials;
provides Ukrainian translation of original ("high") reasoning obtained with DeepSeek-V3;
adds “low” reasoning in Ukrainian by DeepSeek-R1 based on original reasoning and task statements;
ships… See the full description on the dataset page: https://huggingface.co/datasets/anon-researcher-ua/ua-codeforces-cots-open-r1.ualpaca-gpt4
Dataset Card for "alpaca-gpt4-cleaned"
This dataset contains Ukrainian Instruction-Following translated by facebook/nllb-200-3.3B
The dataset was originaly shared in this repository: https://github.com/tloen/alpaca-lora
Licensing Information
The dataset is available under the Creative Commons NonCommercial (CC BY-NC 4.0).
kobza-cleaned-ua
kobza-cleaned-ua
Cleaned Ukrainian-language subset of Goader/kobza dataset with Russian content filtered out.
Dataset Details
Dataset Description
This dataset is a cleaned and filtered version of the Goader/kobza corpus, removing Russian language content to create a pure Ukrainian language dataset suitable for training language models.
The original kobza dataset contains ~60B tokens across 97 million documents. This cleaned version maintains ~59B tokens (98.5%… See the full description on the dataset page: https://huggingface.co/datasets/podarok/kobza-cleaned-ua.UAE-corptax-training
Dataset Card for UAE Corporate Tax Q&A Dataset
Dataset Summary
This dataset contains 1,283 instruction-response pairs covering UAE Corporate Tax regulations from 2022-2025. Built from several official sources. Each response includes proper legal citations.
Perfect for fine-tuning LLMs for UAE tax advisory, building RAG systems, or training tax compliance tools. This is for educational purpose only.
Supported Tasks
Instruction Following: Train models to answer… See the full description on the dataset page: https://huggingface.co/datasets/vikramlingam/UAE-corptax-training.ua_cbt_stories
Dataset Card for UA-CBT Stories
This dataset was generated in the context of Eval-UA-tion 1.0 benchmark for evaluating Ukrainian language models (paper, thesis). It contains the generated and manually corrected stories used for the UA-CBT (Ukrainian Children's Book Test) task.
The dataset contains Ukrainian-language stories, LLM-generated in multiple steps and then manually corrected (or marked as unusable if fixing them was too hard). For each story, the original LLM prompt, all… See the full description on the dataset page: https://huggingface.co/datasets/shamotskyi/ua_cbt_stories.UA-Safety-Align-Sample
UA-Safety-Align: Ukrainian Red Teaming & Safety Dataset (Sample)
Developed by DavidLab
Contact: founder@davidlab.techStatus: Sample (50 examples). Full dataset (4,800+ examples) available for commercial licensing.
🚀 Dataset Summary
UA-Safety-Align is a specialized dataset designed for Red Teaming, Safety Alignment, and Robustness Testing of Large Language Models (LLMs) specifically in the Ukrainian language context.
While most safety datasets focus on English… See the full description on the dataset page: https://huggingface.co/datasets/alexshynkarenk0/UA-Safety-Align-Sample.uae-laws
Dataset Card for UAE-Laws
This dataset is a collection of information about the laws and regulations in the United Arab Emirates.
It covers different areas of law like:
economy and business
family and community
finance and banking
industry and technical standardisation
justice and juiciary, labour
residency and leberal professions
security and safety
tax
Dataset Sources
United Arab Emirates Legislations
Dataset Structure
The ./uae-laws.csv… See the full description on the dataset page: https://huggingface.co/datasets/obadabaq/uae-laws.UAlpaca2.0
UAlpaca 2.0
uae-adab-tutor-600
UAE Adab Tutor 600
This is the 600-conversation supervised fine-tuning dataset used for
adarshrajesh/uae-adab-tutor-qwen3-4b.
Release version: exact-silver v1 Complete-600.
Behavior spec
Across a pressured multi-turn lesson, the tutor should teach the academic
content accurately, correct the specific work without humiliating the learner,
protect learner authorship and assessment integrity, allow respectful
evidence-based disagreement with adults, avoid religious… See the full description on the dataset page: https://huggingface.co/datasets/adarshrajesh/uae-adab-tutor-600.ua-council-decisions
Ukrainian Municipal Council Decisions — Masthead Identity Extraction
Structured-extraction dataset of 1,075 Ukrainian municipal council decisions (рішення) from
43 local councils (громади / ради) — balanced to exactly 25 decisions per council, each paired with the five identity fields that appear in the
document masthead. The task: given the full text of a single decision, extract its masthead identity.
These are public government records. All personal names in the data are… See the full description on the dataset page: https://huggingface.co/datasets/oshyshatskyi/ua-council-decisions.ua-code-bench
LLM Code Generation Benchmark for Ukrainian language
Preprint: https://arxiv.org/pdf/2511.05040
Updates
17/10/2025: paper presented at "Informatics. Culture. Technology" conference;
18/09/2025: added data preparation and evaluation notebooks (check notebooks readme first);
17/09/2025: updated result chart; added gpt-5, gpt-oss, and grok-4 evaluations.
Thousands of programming tasks in Ukrainian language combined with graded Python solutions (code… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/ua-code-bench.ua-toxic-light
Ukrainian Style Chat Mix
Chat-format dataset for Ukrainian style adaptation.
Splits
train: 5820
validation: 90
test: 90
Schema
Each row has:
messages: list of chat turns (role, content)
source: source dataset id
optional style_toxic: 0/1 style marker
Notes
Intended for controlled style tuning.
Keep style data as a minority share during model training.
ua-code-bench
LLM Code Generation Benchmark for Ukrainian language
Preprint: https://arxiv.org/pdf/2511.05040
Updates
18/09/2025: added data preparation and evaluation notebooks (check notebooks readme first);
17/09/2025: updated result chart; added gpt-5, gpt-oss, and grok-4 evaluations;
17/09/2025: paper - end of September, stay tuned;
Thousands of programming tasks in Ukrainian language combined with graded Python solutions (code + reasoning) by leading LLMs (DeepSeek… See the full description on the dataset page: https://huggingface.co/datasets/anon-researcher-ua/ua-code-bench.
