datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
uzbek_homonym_affixes
Uzbek Homonym Affixes Dataset
Dataset link on Hugging Face
📖 Description
This dataset contains Uzbek homonym affixes (omonim qo‘shimchalar) with their occurrences in different parts of speech.The dataset is designed to support Uzbek NLP research, especially in the fields of:
Morphological analysis
Part-of-speech tagging
Word sense disambiguation
Computational linguistics
Each row represents an affix and its possible usage across multiple word classes.… See the full description on the dataset page: https://huggingface.co/datasets/dasturbek/uzbek_homonym_affixes.uzbek-idioms
Uzbek Idioms
Structured dataset of Uzbek idiomatic expressions extracted from O'zbek tili frazeologik lug'ati (2022). Made it for fine-tuning since no machine-readable version of this dictionary existed, so here's one.
4,408 entries. Each row is one meaning of one idiom.
Fields
Field
Description
idiom
The idiomatic expression
syntax_markers
Argument slots: kim?, nima? kimga?, etc.
meaning_uz
Definition in Uzbek
meaning_number
For polysemous… See the full description on the dataset page: https://huggingface.co/datasets/bizb0630/uzbek-idioms.translated_uzbek_dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Xondamir/translated_uzbek_dataset.proverb_Uzbek
🇺🇿 Uzbek Proverbs Dataset (UZBEKPROVERBS-8.5K)
📌 Description
Proverbs are culturally salient figurative expressions, yet Uzbek still lacks standardized NLP resources for their automatic identification. This dataset introduces UzbekProverbs-8.5K, a curated, machine-readable proverb resource for Uzbek, derived from a source inventory of 8,514 entries.
The resource contains:
8,477 unique proverb strings
2,998 glossed entries
70 normalized thematic groups
A structured… See the full description on the dataset page: https://huggingface.co/datasets/ruhilloalaev/proverb_Uzbek.
