CoolFace
Datasetpublic

bevaya/pubmed-ocr

PubMed-OCR: PMC Open Access OCR Annotations PubMed-OCR is an OCR-centric corpus of scientific articles derived from PubMed Central Open Access PDFs. Each page is rendered to an image and annotated with Google Cloud Vision OCR, released in a compact JSON schema with word-, line-, and paragraph-level bounding boxes. Scale (release): 209.5K articles ~1.5M pages ~1.3B words (OCR tokens) This dataset is intended to support layout-aware modeling, coordinate-grounded QA, and… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/pubmed-ocr.

sourceHugging Faceotherupdated 8mo agoView on Hugging Face
72likes3.2kdownloads
13 commits on main
d03682f8mo ago

Update README.md

Hunter Heidenreich
e9dae908mo ago

Update paper link and citation info (#2)

Hunter Heidenreich, nielsr
a9515c18mo ago

Update README.md

Hunter Heidenreich
540f3b08mo ago

Update README.md

Hunter Heidenreich
8580a308mo ago

Add dataset card

Hunter Heidenreich
1bf0f7c8mo ago

Add files using upload-large-folder tool

Hunter Heidenreich
6c2151d8mo ago

Add files using upload-large-folder tool

Hunter Heidenreich
0d85ef48mo ago

Add files using upload-large-folder tool

Hunter Heidenreich
93beda48mo ago

Add files using upload-large-folder tool

Hunter Heidenreich
e75ec428mo ago

Add files using upload-large-folder tool

Hunter Heidenreich
f81b0b38mo ago

Add files using upload-large-folder tool

Hunter Heidenreich
7ac65238mo ago

Add files using upload-large-folder tool

Hunter Heidenreich
9f38aa98mo ago

initial commit

Hunter Heidenreich