CoolFace
Datasetpublic

tahrirchi/uz-books-v2

Dataset Card for UzBooks V2 Dataset Summary UzBooks V2 is an improved version of the UzBooks book corpus for Uzbek language. It contains nearly 40,000 books in two splits: Split Description Examples lat Fully Latin-transliterated version 38,339 cyr Fully Cyrillic-transliterated version 38,339 What's New in V2? OCR Engine Upgrade: Switched from Tesseract → Google Cloud Vision OCR Cleaner Text: Google OCR produces far fewer… See the full description on the dataset page: https://huggingface.co/datasets/tahrirchi/uz-books-v2.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
5likes277downloads
settings

This repository belongs to tahrirchi on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameuz-books-v2
visibilitypublic
licencemit
gatedno
ownertahrirchi
Account settings
tahrirchi/uz-books-v2 · CoolFace