CoolFace
Datasetpublic

murodbek/uz-books

Dataset Card for BookCorpus Dataset Summary In an effort to democratize research on low-resource languages, we release UzBooks dataset, a cleaned book corpus consisting of nearly 40000 books in Uzbek Language divided into two branches: "original" and "lat," representing the OCRed (Latin and Cyrillic) and fully Latin versions of the texts, respectively. Please refer to our blogpost and paper (Coming soon!) for further details. To load and use dataset, run this… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/uz-books.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
23likes310downloads

murodbek/uz-books · main · files are served by the source, never re-hosted here