CoolFace
Datasetpublicgated

voidful/tw-ocr

tw-ocr Filtered OCR/document parsing records for Traditional Chinese OCR experiments. The public Hugging Face upload is parquet-only for dataset rows. Every parquet row embeds non-empty image bytes and is written with Hugging Face datasets Image feature metadata. The physical parquet storage is the standard Image struct ({bytes, path}) with image.path = null; when loaded through datasets, the image column is an Image feature instead of a generic dict column. The source path is… See the full description on the dataset page: https://huggingface.co/datasets/voidful/tw-ocr.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes6downloads
.gitattributesDownload Raw Back to root

This repository is gated, so its file contents are only served once you have accepted the publisher's terms at Hugging Face. Open it at the source above.