datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-papers-by-subject
arXiv Papers by Subject
A reorganised version of the nick007x/arxiv-papers dataset, partitioned by subject code, year, and month for efficient selective access.
Dataset Description
This dataset contains metadata for over 2.5 million arXiv papers, organised into a hierarchical directory structure that allows users to download only the specific subjects and time periods they need, rather than the entire dataset.
Motivation
The original… See the full description on the dataset page: https://huggingface.co/datasets/permutans/arxiv-papers-by-subject.research-papers
research-papers Dataset
Overview
The Research Papers Dataset is a collection of academic research documents categorized by their primary research topic.
This dataset is designed for tasks such as model finetuning, document classification, optical character recognition (OCR) testing and multimodal document understanding (Feel free to use it however you see fit!).
Curated by: tegridy
Language: English
Format: PDF | MD
Repo Structure
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/tegridydev/research-papers.daily-papers-embeddingsarxiv-ai-ml-100k-papers
license: other
tags:
- arxiv
- ocr
- machine-learning
---
# obswork/arxiv-ai-ml-100k
A 99,999-paper stratified subset of
[`Rendra8631/arxiv-papers`](https://huggingface.co/datasets/Rendra8631/arxiv-papers)
at revision `a2c6afb51332d2744b46308df6917697582f8cd4`, filtered to the primary
subjects `cs.AI`, `cs.CV`, `cs.LG`, and `stat.ML`. Only papers submitted in `2023`-`2025` are included.
This dataset is a build artifact of the OCR… See the full description on the dataset page: https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k-papers.Stocks-Daily-Price
Stocks Daily Price
This dataset includes daily price data for various stocks.
25,986,919 rows over 7,764 symbols, 8 columns, covering 1962-01-02 to 2026-08-05. Refreshed monthly.
Strategies Built on This Data
2,401 papers in the Papers With Backtest catalogue declare this dataset as an input. 2,238 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.35, and 45% clear a t-statistic of 1.96 on their own sample… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Stocks-Daily-Price.arxiv-papersscientific-papers
Scientific Papers - Raw Full Text
~57 million scientific papers with full text, extracted from multiple large-scale academic paper collections. This dataset provides raw full text suitable for pre-training, fine-tuning, or building search indices over scientific literature.
Subsets
Subset
Papers
Size
Source
papers-2
~18.5M
~358 GB
S2ORC papers collection (untitled subset)
papers-3
~27.4M
~198 GB
S2ORC scientific-papers collection
pes2o
~8.2M
~106 GB… See the full description on the dataset page: https://huggingface.co/datasets/scientifi-papers/scientific-papers.Universe-Daily-Price
Universe Daily Price
The index of every symbol carried by the public daily price datasets, with the repository that holds it.
21,693 rows over 20,895 symbols, 4 columns. Updated by Papers With Backtest.
Why It Matters
This is the lookup table the other price datasets need:
Routing: A symbol on its own does not say which file holds it. repo_id answers that in one join, so a strategy that mixes equities, futures and rates loads from the right place without… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Universe-Daily-Price.scientific-papersML-ArXiv-PapersThis dataset contains the subset of ArXiv papers with the "cs.LG" tag to indicate the paper is about Machine Learning.
The core dataset is filtered from the full ArXiv dataset hosted on Kaggle: https://www.kaggle.com/datasets/Cornell-University/arxiv. The original dataset contains roughly 2 million papers. This dataset contains roughly 100,000 papers following the category filtering.
The dataset is maintained by with requests to the ArXiv API.
The current iteration of the dataset only contains… See the full description on the dataset page: https://huggingface.co/datasets/CShorten/ML-ArXiv-Papers.arxiv_papers
ArXiv Papers
Description
ArXiv is an online open-access repository of over 2.4 million scholarly papers covering fields such as computer science, mathematics, physics, quantitative biology, economics, and more.
When uploading papers, authors can choose from a variety of licenses.
This dataset includes text from all papers uploaded under CC BY, CC BY-SA, and CC0 licenses through a three-step pipeline:
first, the latex source files for openly licensed papers were… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_papers.scientific_papersScientific papers datasets contains two sets of long and structured documents.
The datasets are obtained from ArXiv and PubMed OpenAccess repositories.
Both "arxiv" and "pubmed" have two features:
- article: the body of the document, pagragraphs seperated by "/n".
- abstract: the abstract of the document, pagragraphs seperated by "/n".
- section_names: titles of sections, seperated by "/n".Forex-Daily-Price
Forex Daily Price
This dataset includes daily price data for various FX pairs.
437,430 rows over 166 symbols, 7 columns, covering 1970-01-04 to 2026-07-31. Refreshed monthly.
Strategies Built on This Data
292 papers in the Papers With Backtest catalogue declare this dataset as an input. 260 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.21, and 34% clear a t-statistic of 1.96 on their own sample, against 48%… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Forex-Daily-Price.daily-papersIndices-Daily-Price
Indices Daily Price
This dataset includes daily price data for various indices.
815,441 rows over 113 symbols, 8 columns, covering 1927-12-30 to 2026-08-03. Refreshed monthly.
Strategies Built on This Data
1,324 papers in the Papers With Backtest catalogue declare this dataset as an input. 1,226 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.49, and 62% clear a t-statistic of 1.96 on their own sample, against… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Indices-Daily-Price.daily-papers-statsPapersarxiv_papers_filtered
ArXiv Papers
Description
ArXiv is an online open-access repository of over 2.4 million scholarly papers covering fields such as computer science, mathematics, physics, quantitative biology, economics, and more.
When uploading papers, authors can choose from a variety of licenses.
This dataset includes text from all papers uploaded under CC BY, CC BY-SA, and CC0 licenses through a three-step pipeline:
first, the latex source files for openly licensed papers were… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_papers_filtered.arxiv_papers_0_50_19subset_arxiv_papers_with_embeddingsThis dataset is a curated subset of the original arXiv dataset, each entry enriched with a 256-dimensional embedding vector. The embeddings are generated using OpenAI's "text-embedding-3-small" model. For each data point, the embedding is created by concatenating the text of the title, author(s), and abstract into a single string, which is then processed by the embedding model. This approach captures the semantic essence of each document, facilitating tasks such as similarity search… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/subset_arxiv_papers_with_embeddings.papercli-papers-acl0_50_arxiv_papers_1742908384paper-sim-n512-traces
paper-sim-n512-traces
Rollout traces for the n=512 simulated bottle-in-bin evaluation reported in
Table I of the WARP-RM CoRL rebuttal.
Two arms, 512 paired scenes each:
directory prefix
arm
bottles/scene
throughput
all-6
fullhz_vanilla_sh00..15
Vanilla BC (100% of data)
3.885
237/hr
9.4%
fullhz_paperwarp512_sh00..15
WARP-BC (31.5% kept)
4.533
290/hr
25.0%
512 traces per arm = 16 shards x 32 worlds. Each qpos_trace_NNN_sSEED.npz
holds the full qpos trajectory… See the full description on the dataset page: https://huggingface.co/datasets/uynitsuj/paper-sim-n512-traces.0_50_arxiv_papers_1742899358arxiv_papers_0_50_22ETFs-Daily-Price
ETFs Daily Price
This dataset includes daily price data for various ETFs.
10,813,800 rows over 5,802 symbols, 8 columns, covering 1978-01-03 to 2026-08-03. Refreshed monthly.
Strategies Built on This Data
502 papers in the Papers With Backtest catalogue declare this dataset as an input. 476 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.55, and 70% clear a t-statistic of 1.96 on their own sample, against 48%… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/ETFs-Daily-Price.top-conference-papers
Top Conference Papers (2024-2026)
Paper metadata from 10 top computer-science conferences (ML / NLP / IR / Data Mining / RecSys), covering 2024-2026. 38,886 papers in total.
Structure
config = conference: ACL, CIKM, ICLR, ICML, KDD, NeurIPS, RecSys, SIGIR, WSDM, WWW
split = year: 2024, 2025, and 2026 where available
from datasets import load_dataset
ds = load_dataset("yufan/top-conference-papers", "ICLR") # one conference
ds["2024"]… See the full description on the dataset page: https://huggingface.co/datasets/yufan/top-conference-papers.arxiv_papers_50_100_19Stocks-Quarterly-BalanceSheet
Stocks Quarterly Balance Sheet
This dataset includes quarterly balance sheet data for various stocks.
342,800 rows over 6,893 symbols, 39 columns, covering 1983-06-30 to 2026-06-30. Refreshed monthly.
Strategies Built on This Data
1,067 papers in the Papers With Backtest catalogue declare this dataset as an input. 1,012 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.27, and 37% clear a t-statistic of 1.96 on… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Stocks-Quarterly-BalanceSheet.0_50_arxiv_papers_1742906667
