CoolFace
Datasetpublic

nielsr/arxiv-chandra-ocr-250-20260401-l40sx1

arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-250-20260401-l40sx1 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-250-20260401-l40sx1 Source paper IDs in input list: 27,584 Processed IDs recorded in state/processed_ids.txt: 250 Successes: 250 Partial successes: 0 Errors: 0 Next shard index: 25 Updated at: 2026-04-01T16:24:48.639101+00:00… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-250-20260401-l40sx1.

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes19downloads

nielsr/arxiv-chandra-ocr-250-20260401-l40sx1 · main · files are served by the source, never re-hosted here