CoolFace
Datasetpublic

p4ulbr4dl3y/dogovors-unlimited-ocr

Dogovors Unlimited-OCR Dataset OCR/document-layout dataset prepared for fine-tuning baidu/Unlimited-OCR. Files train.jsonl contains one JSON object per document. images/ contains the page images referenced by relative path. JSONL Schema { "images": [ "images/doc_001_page_001.jpg", "images/doc_001_page_002.jpg" ], "question": "Multi page parsing.", "answer": "<PAGE><|det|>title [400, 60, 660, 73]<|/det|>... <PAGE><|det|>text [100… See the full description on the dataset page: https://huggingface.co/datasets/p4ulbr4dl3y/dogovors-unlimited-ocr.

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes17downloads

p4ulbr4dl3y/dogovors-unlimited-ocr · main · files are served by the source, never re-hosted here