CoolFace
Datasetpublic

tahrirchi/uz-books-v2

Dataset Card for UzBooks V2 Dataset Summary UzBooks V2 is an improved version of the UzBooks book corpus for Uzbek language. It contains nearly 40,000 books in two splits: Split Description Examples lat Fully Latin-transliterated version 38,339 cyr Fully Cyrillic-transliterated version 38,339 What's New in V2? OCR Engine Upgrade: Switched from Tesseract → Google Cloud Vision OCR Cleaner Text: Google OCR produces far fewer… See the full description on the dataset page: https://huggingface.co/datasets/tahrirchi/uz-books-v2.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
5likes247downloads
Dataset Card

Dataset Card for UzBooks V2

Dataset Summary

UzBooks V2 is an improved version of the UzBooks book corpus for Uzbek language. It contains nearly 40,000 books in two splits:

SplitDescriptionExamples
latFully Latin-transliterated version38,339
cyrFully Cyrillic-transliterated version38,339

What's New in V2?

  • —OCR Engine Upgrade: Switched from Tesseract → Google Cloud Vision OCR
  • —Cleaner Text: Google OCR produces far fewer recognition errors, especially for mixed-script content
  • —Same Structure & Size: Maintains compatibility with v1 — same splits, same number of examples

Usage

python
from datasets import load_dataset

uz_books2 = load_dataset("tahrirchi/uz-books-v2")

# Access Latin version
print(uz_books2["lat"][0]["text"])

# Access Cyrillic version
print(uz_books2["lat"][0]["text"])

Data Fields

FieldTypeDescription
textstringFull text content of the book

Dataset Creation

Books were collected from various public sources and processed using Google Cloud Vision OCR, which delivers substantially better accuracy than Tesseract for Uzbek text — particularly in handling the coexistence of Latin and Cyrillic scripts. Then, lat and cyr splits were generated using curated transliteration scripts.

Citation

bibtex
@online{Mamasaidov2024UzBooksV2,
    author    = {Mukhammadsaid Mamasaidov and Abror Shopulatov},
    title     = {UzBooks V2 dataset},
    year      = {2026},
    url       = {https://huggingface.co/datasets/tahrirchi/uz-books-v2}
}

Contacts

We believe that this work will enable and inspire all enthusiasts around the world to open the hidden beauty of low-resource languages, in particular Uzbek.

For questions or issues:

  • —m.mamasaidov@tahrirchi.uz
  • —a.shopolatov@tahrirchi.uz