Reza2kn/persian-ocr-bench-submitted10-bbox-crops
Persian OCR benchmark — selected submitted bbox crops This dataset contains the non-empty OCR bboxes from the ten explicitly selected submitted pages in persian_ocr_bench_bbox_review. Each row is one PNG crop. gold_text is the current editable OCR content from the live Argilla bbox field (content_text). Geometry is stored both as source page pixels and as percentages of the source page. The original record ID, external ID, bbox ID, source URL, and SHA-256 hashes are included for… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-bench-submitted10-bbox-crops.
Persian OCR benchmark — selected submitted bbox crops
This dataset contains the non-empty OCR bboxes from the ten explicitly selected submitted pages in persian_ocr_bench_bbox_review.
Each row is one PNG crop. gold_text is the current editable OCR content from the live Argilla bbox field (content_text). Geometry is stored both as source page pixels and as percentages of the source page. The original record ID, external ID, bbox ID, source URL, and SHA-256 hashes are included for exact benchmark joins and provenance.
The six intentionally empty text bboxes on image_067 were excluded, as requested. No Argilla records were modified by this export.
The primary image-folder view intentionally exposes only two columns: image followed by gold_text. Full bbox coordinates, IDs, source record references, and hashes are retained in provenance/metadata_full.jsonl.
