reapxdev/arxiv-papers-scraper
arXiv Papers Scraper Search arXiv and export papers with full abstracts, author lists, subject categories, DOIs, journal references and PDF links. Filter by subject class, keyword, author, affiliation or date window. Rows in this dataset 21,722 Fields 26 Collector runs behind it 92 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/arxiv-papers-scraper/ — 10,624 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/arxiv-papers-scraper.

arXiv Papers Scraper
Search arXiv and export papers with full abstracts, author lists, subject categories, DOIs, journal references and PDF links. Filter by subject class, keyword, author, affiliation or date window.
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a template over a keyword list: a row exists because a run observed it.
A browsable presentation of the subset that carries an addressable arxivId is published as 10,624 entity pages at https://reapx.dev/data/arxiv-papers-scraper/, one page per entity. This dataset is the larger of the two — rows whose payload has no field that can address a page are here and are not published there.
Provenance
Each row carries _run_id and _dataset_id, naming the collector run that produced it, so any row can be traced back to the run that observed it. Rows observed by more than one run are deduplicated on content; 0 duplicate observations were collapsed.
Files
arxiv-papers-scraper.jsonl— one JSON object per row, the canonical formarxiv-papers-scraper.csv— the same rows flattened; nested values are JSON-encoded within their cell so they round-tripdataset.json— schema.orgDatasetmetadata
Loading it
from datasets import load_dataset
ds = load_dataset("reapxdev/arxiv-papers-scraper", split="train")A sample of the published entities
- Calculation of prompt diphoton production cross sections at Tevatron and LHC energies
- Inference for Low-Rank Models
- Tractable Control for Autoregressive Language Generation
- Outsourced diffusion sampling: Efficient posterior inference in latent spaces of generative models
- Enginuity: Building an Open Multi-Domain Dataset of Complex Engineering Diagrams
- ACPO: Anchor-Constrained Perceptual Optimization for Diffusion Models with No-Reference Quality Guidance
- Relevant Walk Search for Explaining Graph Neural Networks
- 2606.12836v1
- Revising RVL-CDIP: Quantifying Errors and Test-Train Overlap
- Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG
- 2607.21377v1
- Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses
Related
- All sources: <https://reapx.dev/data/> · machine-readable index: <https://reapx.dev/llms.txt>
- The collector is a public Apify Actor; agents reach it through <https://mcp.apify.com>
Licence
Collected from public sources. This metadata and the published pages are CC BY 4.0; the underlying records remain under the terms of their originating source.
