bevaya/pubmed-ocr
PubMed-OCR: PMC Open Access OCR Annotations PubMed-OCR is an OCR-centric corpus of scientific articles derived from PubMed Central Open Access PDFs. Each page is rendered to an image and annotated with Google Cloud Vision OCR, released in a compact JSON schema with word-, line-, and paragraph-level bounding boxes. Scale (release): 209.5K articles ~1.5M pages ~1.3B words (OCR tokens) This dataset is intended to support layout-aware modeling, coordinate-grounded QA, and… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/pubmed-ocr.
This repository belongs to bevaya on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
