nielsr/arxiv-chandra-ocr-smoke-20260328-tokenfix
arXiv OCR with Chandra OCR 2 This dataset stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-smoke-20260328-tokenfix Source paper IDs in input list: 2 Processed IDs recorded in state/processed_ids.txt: 2 Successes: 2 Partial successes: 0 Errors: 0 Next shard index: 2 Updated at: 2026-03-28T15:22:09.273986+00:00 Files data/part-*.jsonl.gz: OCR result shards, one JSON object per paper… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-smoke-20260328-tokenfix.
012
