CoolFace
Datasetpublic

nielsr/arxiv-chandra-ocr-250-20260401-l40sx1

arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-250-20260401-l40sx1 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-250-20260401-l40sx1 Source paper IDs in input list: 27,584 Processed IDs recorded in state/processed_ids.txt: 250 Successes: 250 Partial successes: 0 Errors: 0 Next shard index: 25 Updated at: 2026-04-01T16:24:48.639101+00:00… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-250-20260401-l40sx1.

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes19downloads
settings

This repository belongs to nielsr on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namearxiv-chandra-ocr-250-20260401-l40sx1
visibilitypublic
licencenot set
gatedno
ownernielsr
Account settings
nielsr/arxiv-chandra-ocr-250-20260401-l40sx1 · CoolFace