gt-free-ocr-metrics/omnidocbench-render-compare-parquet
OmniDocBench Render-and-Compare — Parquet Edition Parquet-shard repackaging of gt-free-ocr-metrics/omnidocbench-render-compare. Overview The pipeline processes each page of OmniDocBench through a Qwen3.5-122B-A10B OCR model, renders the structured output back to a PNG via HTML (reconstructed), and compares it against the original page scan (masked_original) using reference-free visual metrics. Five OCR extraction variants are provided, each targeting a different… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-render-compare-parquet.
chore: anonymize croissant metadata
chore: remove internal parquet-conversion rationale from README
chore: add render_compare_dataset_croissant.json — anonymized, correct Qwen model URL
chore: remove old croissant.json (replaced by render_compare_dataset_croissant.json)
docs: drop ceiling reference from rai:dataSocialImpact
docs: sanitize rai:dataLimitations and rai:dataBiases
docs: sanitize Limitations & Biases section
docs: clarify image-regions masking phrasing
add enriched croissant.json with RAI fields and sc:ImageObject types
docs: rich README with schema, usage examples, RAI fields
upload 64 parquet shards (4,610 pages × 5 variants)
add dataset card
initial commit
