himalaya-ai/ocr-document-processing-eval
ocr_document_processing_eval Document-processing OCR proxy for digitization, KYC-like numeric fields, and RAG extraction checks. Repo: himalaya-ai/ocr-document-processing-eval Task: document_processing_ocr Main raw file: *.ocr.jsonl with image, ocr, source_repo, and language/provenance columns. Optional fine-tuning/eval file: *.sharegpt.json with messages and images. Core Columns id: unique sample identifier image: relative path to the image file ocr:… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/ocr-document-processing-eval.
ocrdocumentprocessing_eval
Document-processing OCR proxy for digitization, KYC-like numeric fields, and RAG extraction checks.
- Repo:
himalaya-ai/ocr-document-processing-eval - Task:
document_processing_ocr - Main raw file:
*.ocr.jsonlwithimage,ocr,source_repo, and language/provenance columns. - Optional fine-tuning/eval file:
*.sharegpt.jsonwithmessagesandimages.
Core Columns
id: unique sample identifierimage: relative path to the image fileocr: ground-truth text label
Source Mix
indic_vision_bench_deva_ocrocr/test: 40%nayana_bench_deva_documentsdefault/hi: 20%nayana_bench_deva_documentsdefault/mr: 15%devanagari_digits_mixeddefault/train: 15%hindi_handwritten_word_ocrdefault/test: 10%
Notes
Generated by scripts/sample_ocr_eval_sets.py from the GLM fine-tuning workspace. If a source only exposes a train split, keep the deterministic held-out row ids out of SFT/training runs.
