CoolFace
Datasetpublic

nielsr/arxiv-chandra-ocr-smoke-20260328-tokenfix

arXiv OCR with Chandra OCR 2 This dataset stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-smoke-20260328-tokenfix Source paper IDs in input list: 2 Processed IDs recorded in state/processed_ids.txt: 2 Successes: 2 Partial successes: 0 Errors: 0 Next shard index: 2 Updated at: 2026-03-28T15:22:09.273986+00:00 Files data/part-*.jsonl.gz: OCR result shards, one JSON object per paper… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-smoke-20260328-tokenfix.

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes12downloads

nielsr/arxiv-chandra-ocr-smoke-20260328-tokenfix · main · files are served by the source, never re-hosted here