tahrirchi/uz-books-v2
Dataset Card for UzBooks V2 Dataset Summary UzBooks V2 is an improved version of the UzBooks book corpus for Uzbek language. It contains nearly 40,000 books in two splits: Split Description Examples lat Fully Latin-transliterated version 38,339 cyr Fully Cyrillic-transliterated version 38,339 What's New in V2? OCR Engine Upgrade: Switched from Tesseract → Google Cloud Vision OCR Cleaner Text: Google OCR produces far fewer… See the full description on the dataset page: https://huggingface.co/datasets/tahrirchi/uz-books-v2.
updating number of books
uploading lat and cyr splits
uploading lat and cyr splits (part 00001-of-00002)
uploading lat and cyr splits (part 00000-of-00002)
adding more info to the README
fixed issue with
Upload dataset
initial commit
