papers
Datasets
All datasets matching “papers”arxiv-papers-by-subject
arXiv Papers by Subject
A reorganised version of the nick007x/arxiv-papers dataset, partitioned by subject code, year, and month for efficient selective access.
Dataset Description
This dataset contains metadata for over 2.5 million arXiv papers, organised into a hierarchical directory structure that allows users to download only the specific subjects and time periods they need, rather than the entire dataset.
Motivation
The original… See the full description on the dataset page: https://huggingface.co/datasets/permutans/arxiv-papers-by-subject.research-papers
research-papers Dataset
Overview
The Research Papers Dataset is a collection of academic research documents categorized by their primary research topic.
This dataset is designed for tasks such as model finetuning, document classification, optical character recognition (OCR) testing and multimodal document understanding (Feel free to use it however you see fit!).
Curated by: tegridy
Language: English
Format: PDF | MD
Repo Structure
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/tegridydev/research-papers.daily-papers-embeddingsarxiv-ai-ml-100k-papers
license: other
tags:
- arxiv
- ocr
- machine-learning
---
# obswork/arxiv-ai-ml-100k
A 99,999-paper stratified subset of
[`Rendra8631/arxiv-papers`](https://huggingface.co/datasets/Rendra8631/arxiv-papers)
at revision `a2c6afb51332d2744b46308df6917697582f8cd4`, filtered to the primary
subjects `cs.AI`, `cs.CV`, `cs.LG`, and `stat.ML`. Only papers submitted in `2023`-`2025` are included.
This dataset is a build artifact of the OCR… See the full description on the dataset page: https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k-papers.Stocks-Daily-Price
Stocks Daily Price
This dataset includes daily price data for various stocks.
25,986,919 rows over 7,764 symbols, 8 columns, covering 1962-01-02 to 2026-08-05. Refreshed monthly.
Strategies Built on This Data
2,401 papers in the Papers With Backtest catalogue declare this dataset as an input. 2,238 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.35, and 45% clear a t-statistic of 1.96 on their own sample… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Stocks-Daily-Price.arxiv-papers
