datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
preprocessed_hand_skeletonsuz-lexicon-skeleton
Uzbek frequency lexicon with skeleton index
English · Oʻzbekcha
Derived from tahrirchi/uz-books-v2 (MIT)
and tahrirchi/uz-crawl (Apache-2.0).
1.75 billion word tokens counted: 1.43 B from books, 296 M from news, 25.5 M from Telegram channels.
Licence: CC BY 4.0 — credit this dataset and both sources.
4,892,125 Uzbek word forms with their frequencies, 23,160,605 word pairs, and a
skeleton index that groups the words a keyboard makes indistinguishable. Every word is
written in… See the full description on the dataset page: https://huggingface.co/datasets/Maqsudjonpolatov/uz-lexicon-skeleton.
