voidful/tw-ocr
tw-ocr Filtered OCR/document parsing records for Traditional Chinese OCR experiments. The public Hugging Face upload is parquet-only for dataset rows. Every parquet row embeds non-empty image bytes and is written with Hugging Face datasets Image feature metadata. The physical parquet storage is the standard Image struct ({bytes, path}) with image.path = null; when loaded through datasets, the image column is an Image feature instead of a generic dict column. The source path is… See the full description on the dataset page: https://huggingface.co/datasets/voidful/tw-ocr.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face