datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
uzbek-instruct-llmuzbek-instruct-llm is a corpus of more than 15,000 records. It's made for instruct fine-tuning large language models for Uzbek language. It's mostly translated from other instruct datasets with some extra data added
uzbek-asr-curated-701h
Uzbek ASR Curated Dataset (701 hours)
A curated multi-source Uzbek speech dataset for automatic speech recognition (ASR) training and evaluation.
Dataset Description
Language
Uzbek (Latin script with okina ʻ)
Total utterances
337,920
Total duration
~701 hours
Audio format
16 kHz mono WAV (PCM_16)
Manifest format
NeMo JSONL
Splits
train (94%) / val (3%) / test (3%)
Splits
Split
Utterances
Hours
Train
317,655… See the full description on the dataset page: https://huggingface.co/datasets/uzinfocom-edu-ai/uzbek-asr-curated-701h.uzbek-legal-corpus
Uzbek Legal Corpus (Oʻzbek huquqiy korpusi)
25 cleaned, audited, machine-readable texts of the Republic of Uzbekistan's Constitution, 20 codes, and 4 major laws — in Uzbek (Latin script), sourced from the official National Database of Legislation (lex.uz). Snapshot: July 2026. 7,368 articles.
Built for NLP / LLM / RAG work on Uzbek legal text. Released by Tomaris AI.
Contents
Path
What it is
data/raw/*.txt
Cleaned plain text, one file per code/law… See the full description on the dataset page: https://huggingface.co/datasets/javohirmat/uzbek-legal-corpus.uzbek_ner
Uzbek NER Dataset
About the Dataset
This dataset is created for Named Entity Recognition (NER) in Uzbek texts. The dataset includes named entities from various categories such as persons, places, organizations, dates, and more.
Data Structure
The data is provided in JSON format with the following structure:
{
"LOC": ["Location names"],
"ORG": ["Organization names"],
"PERSON": ["Person names"],
"DATE": ["Date expressions"],
"MONEY":… See the full description on the dataset page: https://huggingface.co/datasets/risqaliyevds/uzbek_ner.uzbek-asr-train-manifests
Uzbek ASR Training Manifests
The exact training, validation and test splits behind
rustam1221/uzbek-asr-gigaam:
974 hours of Uzbek speech drawn from seven public corpora, filtered, text-normalized,
and split by speaker.
No audio is copied. Each row is a pointer — a parquet file plus a row
index in the upstream dataset — and the training dataloader decodes the audio
when the batch is built. That keeps the whole corpus definition at 200 MB
instead of roughly a terabyte of… See the full description on the dataset page: https://huggingface.co/datasets/rustam1221/uzbek-asr-train-manifests.soas-english-uzbek-rag-evaluation
SOAS English-Uzbek Retrieval Pilot
Dataset Summary
This folder documents a bilingual English-Uzbek retrieval evaluation benchmark for culturally grounded RAG systems. The 400-row public pilot release is retrieval-only: it contains questions and source-document targets, but it intentionally excludes answer, context, excerpt, and source-text fields.
This is a pilot benchmark with documented quality flags, template-generated examples, and domain mismatches. The rows… See the full description on the dataset page: https://huggingface.co/datasets/Rajan2026/soas-english-uzbek-rag-evaluation.uzbek_spam_dataset
Uzbek Spam Detection Dataset
A synthetic dataset for training spam detection models for Uzbek language, specifically targeting Telegram-style messages.
Dataset Description
This dataset contains 2000 Uzbek text messages labeled as either "spam" or "normal".
Dataset Summary
Total samples: 2000
Train split: 1800
Test split: 200
Labels: spam, normal
Language: Uzbek (Latin and Cyrillic scripts)
Spam Categories Covered
Aggressive advertising and… See the full description on the dataset page: https://huggingface.co/datasets/sukhrobnurali/uzbek_spam_dataset.uzbek_constitution
Constitution of the Republic of Uzbekistan (Trilingual Dataset)
Dataset Summary
This dataset contains the Constitution of the Republic of Uzbekistan (New Edition, adopted via referendum on April 30, 2023) aligned in three languages:
🇺🇿 Uzbek (Latin script)
🇷🇺 Russian
🇬🇧 English
The data was carefully scraped and processed from the official National Database of Legislation of the Republic of Uzbekistan. It serves as a high-quality parallel corpus for legal NLP… See the full description on the dataset page: https://huggingface.co/datasets/sukhrobnurali/uzbek_constitution.English-Uzbek-Translation-1
🌐 English–Uzbek Translation Dataset
A parallel corpus for English ↔ Uzbek machine translation, curated to support research and development of NLP models for the Uzbek language — one of the most underrepresented Turkic languages in open-source datasets.
📖 Dataset Description
This dataset contains aligned sentence pairs in English and Uzbek, designed for training, fine-tuning, and evaluating neural machine translation (NMT) models. The dataset aims to bridge the resource… See the full description on the dataset page: https://huggingface.co/datasets/ML-Jonibek/English-Uzbek-Translation-1.uzbeklex-uzbek-laws
Lex.uz — “Xavfsizlik va huquq-tartibot muhofazasi” (Partial) Dataset (Uzbek)
Dataset Summary
This dataset contains Uzbek legal texts collected from the lex.uz portal, specifically from the “Xavfsizlik va huquq-tartibot muhofazasi” (Security and law enforcement protection) section.The collection is partial: the full section has not been fully scraped yet, and the dataset may be expanded in future updates.
The dataset is intended for:
domain adaptation / fine-tuning of… See the full description on the dataset page: https://huggingface.co/datasets/Mehriddin1997/lex-uzbek-laws.uzbekdataEnglish-Uzbek-Translationuzbek-asr-full-data
Saidzoda Lab — Gated Research Dataset
Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz).
Dataset contents, provenance, and statistics are not publicly disclosed.
Access is granted manually on request.
customer-support-uzbek-120bEnglish-Uzbek-Translation
