sopheakvoatei/truncated_rendered_kh_tables_suryaocr2
Surya OCR 2 (table) on sopheakvoatei/rendered_khmer_tables Table recognition (mode full) over images in sopheakvoatei/rendered_khmer_tables using Surya OCR 2 (650M, Qwen3.5-based) by Datalab, via the surya-ocr package, run as offline vLLM batch inference on Hugging Face Jobs. Processing Details Source Dataset: sopheakvoatei/rendered_khmer_tables Model: datalab-to/surya-ocr-2 Task: table (table mode full) Input column: image (image) Text column: markdown… See the full description on the dataset page: https://huggingface.co/datasets/sopheakvoatei/truncated_rendered_kh_tables_suryaocr2.
Surya OCR 2 (table) on sopheakvoatei/renderedkhmertables
Table recognition (mode full) over images in sopheakvoatei/rendered_khmer_tables using Surya OCR 2 (650M, Qwen3.5-based) by Datalab, via the `surya-ocr` package, run as offline vLLM batch inference on Hugging Face Jobs.
Processing Details
- Source Dataset: sopheakvoatei/rendered_khmer_tables
- Model: datalab-to/surya-ocr-2
- Task:
table(table modefull) - Input column:
image(image) - Text column:
markdown(flattened, reading-order text per row) - Structured column:
surya_blocks(JSON: per-page blocks with bbox / polygon / label / reading_order / confidence / html) - Split:
train - Samples: 892
- Processed OK: 892 / 892
- Processing time: 43.0 min
- Date: 2026-08-19 04:08 UTC
License note
Surya's code is Apache-2.0, but the model weights use a modified OpenRAIL-M license: free for research, personal use, and startups under $5M funding/revenue, restricted from competitive use against Datalab's API. See the model card.
Dataset Structure
Original columns plus:
markdown: flattened text (OCR), label outline (layout), or table HTML (table)surya_blocks: structured result as a JSON string (one entry per page)inference_info: JSON list tracking models applied to this dataset
Generated with UV Scripts.
