datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-metadata-2020-2026
arXiv Metadata, enriched (2020–2026)
Per-paper metadata for 1,517,185 arXiv papers spanning 2020-01 → 2026-09,
enriched with abstracts, citation counts, and Semantic Scholar identifiers, and
organized as a two-level hierarchy: field of study → year.
Unlike a bare title index, every record here carries the abstract, the
full author list, citation counts, and the Semantic Scholar corpusId,
so you can do retrieval, classification, citation analysis, and corpus building
directly… See the full description on the dataset page: https://huggingface.co/datasets/yufan/arxiv-metadata-2020-2026.arxiv-software-repo-links
arXiv Software Repository Links
A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering.
Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models
Quick Start
from datasets import load_dataset
# Load DOI-to-repo links
links = load_dataset("cometadata/arxiv-software-repo-links", "links")
#… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links.arxiv-chandra-ocr-2-include-images-first50-20260415
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415
Source paper IDs in input list: 27,584
Processed IDs recorded in state/processed_ids.txt: 50
Successes: 50
Partial successes: 0
Errors: 0
Next shard index: 10
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415.ArxivSummarizemathpile_arxiv_subset_tiny
MathPile ArXiv (subset)
Description
This dataset consists of a toy subset of 8834 (5000 training + 3834 testing) TeX files found in the arXiv subset of MathPile, used for testing. You should not use this dataset. Training and testing sets are already split
Source
The data was obtained from the training + validation portion of the arXiv subset of MathPile.
Format
Given as JSONL files of JSON dicts each containing the single key: "text"
Usage… See the full description on the dataset page: https://huggingface.co/datasets/aluncstokes/mathpile_arxiv_subset_tiny.arxiv-author-affiliation-extraction-inference-inputs-metadataarxiv-software-repo-links
arXiv Software Repository Links
A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering.
Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models
Quick Start
from datasets import load_dataset
# Load DOI-to-repo links
links = load_dataset("cometadata/arxiv-software-repo-links", "links")
#… See the full description on the dataset page: https://huggingface.co/datasets/rafidirtiza/arxiv-software-repo-links.arxiv-papers-scraper
arXiv Papers Scraper
Search arXiv and export papers with full abstracts, author lists, subject categories, DOIs, journal references and PDF links. Filter by subject class, keyword, author, affiliation or date window.
Rows in this dataset
21,722
Fields
26
Collector runs behind it
92
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/arxiv-papers-scraper/ — 10,624 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/arxiv-papers-scraper.arxivmath-chunk-summaries
ArXivMath chunk summaries (Qwen3.5-9B, run v2)
Short summaries of each chunk of a long-form math solution, generated offline with
Qwen/Qwen3.5-9B. The summaries are intended as supervision targets for belief /
compaction tokens in a chunked recurrent reasoning model: instead of predicting the raw
next chunk, the model is trained to predict a compressed summary of it.
Source: MathArena/arxivmath-training_outputs
at revision 02002a6d4e39033de27adeb4e4683deeb6f22850. The chunk… See the full description on the dataset page: https://huggingface.co/datasets/BhavyaAI139/arxivmath-chunk-summaries.arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416.arxiv-embeddingsarxiv-cs2021-embeddings-bge-m3arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416.arxiv-chandra-ocr-250-20260401-l40sx1
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-250-20260401-l40sx1
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-250-20260401-l40sx1
Source paper IDs in input list: 27,584
Processed IDs recorded in state/processed_ids.txt: 250
Successes: 250
Partial successes: 0
Errors: 0
Next shard index: 25
Updated at: 2026-04-01T16:24:48.639101+00:00… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-250-20260401-l40sx1.arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1
Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417.roaming-arxiv-ai-ml-20250405-20260405
arxiv-ai-ml-20250405-20260405 (JSON)
Large machine-oriented export of arXiv papers (cs.AI, cs.LG, cs.CL, stat.ML, cs.NE) for the date window documented in the JSON date_range_utc field.
Linked GitHub repository
Canonical workspace: github.com/msandroid/roaming
Human-readable companion: arxiv-ai-ml-20250405-20260405.md in that repo.
Schema: top-level schema: roaming.arxiv_feed.v1 in the JSON.
Source API: https://export.arxiv.org/api/query
arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416.arxiv-chandra-ocr-smoke-20260328-tokenfix
arXiv OCR with Chandra OCR 2
This dataset stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-smoke-20260328-tokenfix
Source paper IDs in input list: 2
Processed IDs recorded in state/processed_ids.txt: 2
Successes: 2
Partial successes: 0
Errors: 0
Next shard index: 2
Updated at: 2026-03-28T15:22:09.273986+00:00
Files
data/part-*.jsonl.gz: OCR result shards, one JSON object per paper… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-smoke-20260328-tokenfix.arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0
Next shard index: 1… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416.arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-20260416.arxiv1m-zeronex⭐ README — ZERONEX SCIENTIFIC CORPUS (1M CLEAN JSON)
(Made by Zeronex — 2025 Edition)
🚀 Overview
This release contains one of the cleanest scientific corpora ever published.
No noise. No XML leftovers. No broken paragraphs.
Every file is fully normalized, token-ready, embedding-ready, and AI-training-ready.
All files are professionally structured JSON, signature-stamped, and extracted from scientific metadata with gold-level cleaning rules.
This drop includes:
1️⃣ The MASSIVE 1,000,000 Sample… See the full description on the dataset page: https://huggingface.co/datasets/YoloMG/arxiv1m-zeronex.arxiv-chandra-ocr-full-20260328-p30
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-full-20260328-p30
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-full-20260328-p30
Source paper IDs in input list: 27,584
Processed IDs recorded in state/processed_ids.txt: 250
Successes: 250
Partial successes: 0
Errors: 0
Next shard index: 25
Updated at: 2026-03-29T01:12:18.345802+00:00… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-full-20260328-p30.arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-v2-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-v2-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-v2-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-v2-20260416.arxiv-classifier-leaderboard-requestsArXiv_Cybersecurity
