CoolFace
Datasetpublic

Maqsudjonpolatov/uz-lexicon-skeleton

Uzbek frequency lexicon with skeleton index English · Oʻzbekcha Derived from tahrirchi/uz-books-v2 (MIT) and tahrirchi/uz-crawl (Apache-2.0). 1.75 billion word tokens counted: 1.43 B from books, 296 M from news, 25.5 M from Telegram channels. Licence: CC BY 4.0 — credit this dataset and both sources. 4,892,125 Uzbek word forms with their frequencies, 23,160,605 word pairs, and a skeleton index that groups the words a keyboard makes indistinguishable. Every word is written in… See the full description on the dataset page: https://huggingface.co/datasets/Maqsudjonpolatov/uz-lexicon-skeleton.

sourceHugging Facecc-by-4.0updated 11d agoView on Hugging Face
2likes48downloads
settings

This repository belongs to Maqsudjonpolatov on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameuz-lexicon-skeleton
visibilitypublic
licencecc-by-4.0
gatedno
ownerMaqsudjonpolatov
Account settings
Maqsudjonpolatov/uz-lexicon-skeleton · CoolFace