datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
uzbek-instruct-llmuzbek-instruct-llm is a corpus of more than 15,000 records. It's made for instruct fine-tuning large language models for Uzbek language. It's mostly translated from other instruct datasets with some extra data added
ua-tg-misc
UA/TG misc — 15 small channel stubs
A bundle of the 15 small Telegram channel exports (10 of them kept non-empty text rows; the rest exported only media-only posts, stripped here) — stubs and partially
scraped channels that did not reach the full 10k-message scrape that the
main corpora (hausmer/ukr-tg-media,
hausmer/ukr-tg-satire)
got. Most are 1–30 text posts (some channels were only reachable briefly,
others export empty media rows).
This corpus contains 67 text posts across… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/ua-tg-misc.CVEs
CVEs — a full-coverage CVE chat dataset
1,625,017 chat conversations covering all 361,190 usable CVEs (1999–2026), built for fine-tuning cybersecurity assistants. Every known CVE in the official CVE List with severity enrichment from NVD (via the fkie-cad community feeds), rendered as English user/assistant conversations with varied phrasings, honest handling of missing data, and a per-CVE 99/1 train/validation split with zero leakage.
The schema matches oi-uae/cyber-security… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/CVEs.uae-laws-irac
UAE Laws Q&A Dataset (IRAC Format)
A high-quality dataset of 9,477 question-answer pairs about UAE laws, formatted in IRAC (Issue, Rule, Application, Conclusion) legal reasoning structure.
Dataset Creation
Source Documents
The dataset was built from a comprehensive collection of UAE legal documents, including:
Federal Decrees and Laws
Cabinet Resolutions
Ministerial Decisions
Civil and Commercial Codes
Labor Law
Traffic Law
And more
Creation Process… See the full description on the dataset page: https://huggingface.co/datasets/SalahALHaismawi/uae-laws-irac.UAE-corptax-training
Dataset Card for UAE Corporate Tax Q&A Dataset
Dataset Summary
This dataset contains 1,283 instruction-response pairs covering UAE Corporate Tax regulations from 2022-2025. Built from several official sources. Each response includes proper legal citations.
Perfect for fine-tuning LLMs for UAE tax advisory, building RAG systems, or training tax compliance tools. This is for educational purpose only.
Supported Tasks
Instruction Following: Train models to answer… See the full description on the dataset page: https://huggingface.co/datasets/vikramlingam/UAE-corptax-training.UA-Safety-Align-Sample
UA-Safety-Align: Ukrainian Red Teaming & Safety Dataset (Sample)
Developed by DavidLab
Contact: founder@davidlab.techStatus: Sample (50 examples). Full dataset (4,800+ examples) available for commercial licensing.
🚀 Dataset Summary
UA-Safety-Align is a specialized dataset designed for Red Teaming, Safety Alignment, and Robustness Testing of Large Language Models (LLMs) specifically in the Ukrainian language context.
While most safety datasets focus on English… See the full description on the dataset page: https://huggingface.co/datasets/alexshynkarenk0/UA-Safety-Align-Sample.UAlpaca2.0
UAlpaca 2.0
uae-adab-tutor-600
UAE Adab Tutor 600
This is the 600-conversation supervised fine-tuning dataset used for
adarshrajesh/uae-adab-tutor-qwen3-4b.
Release version: exact-silver v1 Complete-600.
Behavior spec
Across a pressured multi-turn lesson, the tutor should teach the academic
content accurately, correct the specific work without humiliating the learner,
protect learner authorship and assessment integrity, allow respectful
evidence-based disagreement with adults, avoid religious… See the full description on the dataset page: https://huggingface.co/datasets/adarshrajesh/uae-adab-tutor-600.ua-council-decisions
Ukrainian Municipal Council Decisions — Masthead Identity Extraction
Structured-extraction dataset of 1,075 Ukrainian municipal council decisions (рішення) from
43 local councils (громади / ради) — balanced to exactly 25 decisions per council, each paired with the five identity fields that appear in the
document masthead. The task: given the full text of a single decision, extract its masthead identity.
These are public government records. All personal names in the data are… See the full description on the dataset page: https://huggingface.co/datasets/oshyshatskyi/ua-council-decisions.ua-toxic-light
Ukrainian Style Chat Mix
Chat-format dataset for Ukrainian style adaptation.
Splits
train: 5820
validation: 90
test: 90
Schema
Each row has:
messages: list of chat turns (role, content)
source: source dataset id
optional style_toxic: 0/1 style marker
Notes
Intended for controlled style tuning.
Keep style data as a minority share during model training.
