CoolFace
Datasetpublic

tahrirchi/uz-books-v2

Dataset Card for UzBooks V2 Dataset Summary UzBooks V2 is an improved version of the UzBooks book corpus for Uzbek language. It contains nearly 40,000 books in two splits: Split Description Examples lat Fully Latin-transliterated version 38,339 cyr Fully Cyrillic-transliterated version 38,339 What's New in V2? OCR Engine Upgrade: Switched from Tesseract → Google Cloud Vision OCR Cleaner Text: Google OCR produces far fewer… See the full description on the dataset page: https://huggingface.co/datasets/tahrirchi/uz-books-v2.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
5likes278downloads
8 commits on main
aac9c106mo ago

updating number of books

murodbek
e1678846mo ago

uploading lat and cyr splits

murodbek
9ad84846mo ago

uploading lat and cyr splits (part 00001-of-00002)

murodbek
806befa6mo ago

uploading lat and cyr splits (part 00000-of-00002)

murodbek
785969a6mo ago

adding more info to the README

murodbek
83318703y ago

fixed issue with

murodbek
fe01d483y ago

Upload dataset

murodbek
b0d9d013y ago

initial commit

murodbek