CoolFace
Datasetpublic

christinamaria/cord-v2-ocr

CORD-v2 text-only — receipt OCR text → structured JSON A text-only derivative of CORD-v2 (Consolidated Receipt Dataset; Park et al., 2019), the standard benchmark for document information extraction used by Donut and similar models. The original dataset pairs 1,000 receipt photos with a rich ground-truth schema (~30 field types across 4 groups, including menu line items). This version drops the images and pairs the OCR text of each receipt with its target parse, so text-only… See the full description on the dataset page: https://huggingface.co/datasets/christinamaria/cord-v2-ocr.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes67downloads
settings

This repository belongs to christinamaria on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namecord-v2-ocr
visibilitypublic
licencecc-by-4.0
gatedno
ownerchristinamaria
Account settings
christinamaria/cord-v2-ocr · CoolFace