datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
uzbek_NER
Uzbek NER Gold
Uzbek NER Gold is a token-level named entity recognition dataset for Uzbek. The dataset is distributed as a UTF-8 TSV file and uses BIO tagging for named entities.
Dataset Summary
Dataset ID: uznlp-uz/uzbek_NER
Language: Uzbek (uz)
Rows: 59,569 token rows
Columns: 5
Sentences: 4,176
Split: train
Format: UTF-8 TSV
Data file: Uzbek_NER_Gold.tsv
License: CC BY 4.0
Data Fields
Field
Description
Sentence
Sentence identifier.… See the full description on the dataset page: https://huggingface.co/datasets/uznlp-uz/uzbek_NER.Uzbek_and_Russian_Phishing_dataset
Citation
If you use this work, please cite it as follows:
@inproceedings{saydiev2025uzphishnet,
title={UzPhishNet Model for Phishing Detection},
author={Saydiev, Bektemir and Cui, Xiaohui and Zukaib, Umer},
booktitle={International Conference on Information and Communications Security},
pages={501--512},
year={2025},
organization={Springer}
}
uzbek_sa
Sentiment Analysis Data for the Uzbek Language
Dataset Description:
This dataset contains a sentiment analysis dataset from Kuriyozov et al. (2019).
Data Structure:
The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages.
Citation:
@inproceedings{DBLP:conf/ltconf/KuriyozovMAG19,
author = {Elmurod Kuriyozov and
Sanatbek Matlatipov and
Miguel A. Alonso and
Carlos… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/uzbek_sa.Finance-Curriculum-Edu-Uzbek
Finance Curriculum Edu Uzbek
Dataset Summary
English:Finance Curriculum Edu Uzbek is an expansive, curated Q&A dataset covering the full range of practical and professional finance, with a special focus on the Uzbek language and context. Part of a larger multilingual project, this dataset is designed for language model training, benchmarking, and financial education. It features over 2,200 unique, carefully categorized seed questions mapped to a multi-level finance… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/Finance-Curriculum-Edu-Uzbek.uzbek_homonym_affixes
Uzbek Homonym Affixes Dataset
Dataset link on Hugging Face
📖 Description
This dataset contains Uzbek homonym affixes (omonim qo‘shimchalar) with their occurrences in different parts of speech.The dataset is designed to support Uzbek NLP research, especially in the fields of:
Morphological analysis
Part-of-speech tagging
Word sense disambiguation
Computational linguistics
Each row represents an affix and its possible usage across multiple word classes.… See the full description on the dataset page: https://huggingface.co/datasets/dasturbek/uzbek_homonym_affixes.uzbek-idioms
Uzbek Idioms
Structured dataset of Uzbek idiomatic expressions extracted from O'zbek tili frazeologik lug'ati (2022). Made it for fine-tuning since no machine-readable version of this dictionary existed, so here's one.
4,408 entries. Each row is one meaning of one idiom.
Fields
Field
Description
idiom
The idiomatic expression
syntax_markers
Argument slots: kim?, nima? kimga?, etc.
meaning_uz
Definition in Uzbek
meaning_number
For polysemous… See the full description on the dataset page: https://huggingface.co/datasets/bizb0630/uzbek-idioms.Uzbek-restaurant-domain-sentiment-reviewsuzbek-legal-datasetzero-shot-classification-news-uzbekuzbek_settlements
🏙️ Uzbek Settlements Dataset
The Uzbekistan Settlements Dataset is a structured list of administrative and populated places across Uzbekistan.It contains hierarchical information about regions (viloyatlar), districts (tumanlar), cities, and villages, represented with simple relational fields.
📦 Dataset Summary
Total Records: (depends on coverage)
Language: Uzbek (uz)
Format: Parquet / CSV
Split: train (complete dataset)
This dataset is useful for:
🏘️ Building… See the full description on the dataset page: https://huggingface.co/datasets/databek/uzbek_settlements.proverb_Uzbek
🇺🇿 Uzbek Proverbs Dataset (UZBEKPROVERBS-8.5K)
📌 Description
Proverbs are culturally salient figurative expressions, yet Uzbek still lacks standardized NLP resources for their automatic identification. This dataset introduces UzbekProverbs-8.5K, a curated, machine-readable proverb resource for Uzbek, derived from a source inventory of 8,514 entries.
The resource contains:
8,477 unique proverb strings
2,998 glossed entries
70 normalized thematic groups
A structured… See the full description on the dataset page: https://huggingface.co/datasets/ruhilloalaev/proverb_Uzbek.translated_uzbek_dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Xondamir/translated_uzbek_dataset.russian-uzbek-names-cyrillic
