ShayManor/Labeled-arXiv
Labeled-arXiv Enriched arXiv paper metadata and derived author bibliometrics sourced from OpenAlex. Contains 2.93M papers and 1.72M authors across all arXiv subject areas. Subsets papers — 2.93M rows Column Type Description id string arXiv paper ID (e.g. 0704.0028) submitter string (nullable) Name of the submitting author authors string Raw author string from arXiv metadata title string Paper title comments string (nullable)… See the full description on the dataset page: https://huggingface.co/datasets/ShayManor/Labeled-arXiv.
Labeled-arXiv
Enriched arXiv paper metadata and derived author bibliometrics sourced from OpenAlex. Contains 2.93M papers and 1.72M authors across all arXiv subject areas.
Subsets
papers — 2.93M rows
authors — 1.72M rows
Author-level metrics derived from the papers in this dataset, not global totals.
Note:h_index,works_count, andcited_by_countare scoped to this dataset and do not represent an author's complete publication record.
Usage
from datasets import load_dataset
papers = load_dataset("ShayManor/Labeled-arXiv", "papers", split="train")
authors = load_dataset("ShayManor/Labeled-arXiv", "authors", split="train")Source
Metadata from OpenAlex (Priem et al., 2022), built on top of the arXiv bulk metadata. Author IDs use ORCID URLs or OpenAlex internal identifiers.
License
CC0 1.0 — Public Domain.
