voidful/tw-ocr
tw-ocr Filtered OCR/document parsing records for Traditional Chinese OCR experiments. The public Hugging Face upload is parquet-only for dataset rows. Every parquet row embeds non-empty image bytes and is written with Hugging Face datasets Image feature metadata. The physical parquet storage is the standard Image struct ({bytes, path}) with image.path = null; when loaded through datasets, the image column is an Image feature instead of a generic dict column. The source path is… See the full description on the dataset page: https://huggingface.co/datasets/voidful/tw-ocr.
06
No card is published for this repository, or it could not be fetched from Hugging Face right now.
