Maqsudjonpolatov/uz-lexicon-skeleton
Uzbek frequency lexicon with skeleton index English · Oʻzbekcha Derived from tahrirchi/uz-books-v2 (MIT) and tahrirchi/uz-crawl (Apache-2.0). 1.75 billion word tokens counted: 1.43 B from books, 296 M from news, 25.5 M from Telegram channels. Licence: CC BY 4.0 — credit this dataset and both sources. 4,892,125 Uzbek word forms with their frequencies, 23,160,605 word pairs, and a skeleton index that groups the words a keyboard makes indistinguishable. Every word is written in… See the full description on the dataset page: https://huggingface.co/datasets/Maqsudjonpolatov/uz-lexicon-skeleton.
This repository belongs to Maqsudjonpolatov on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
