datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-papers-by-subject
arXiv Papers by Subject
A reorganised version of the nick007x/arxiv-papers dataset, partitioned by subject code, year, and month for efficient selective access.
Dataset Description
This dataset contains metadata for over 2.5 million arXiv papers, organised into a hierarchical directory structure that allows users to download only the specific subjects and time periods they need, rather than the entire dataset.
Motivation
The original… See the full description on the dataset page: https://huggingface.co/datasets/permutans/arxiv-papers-by-subject.scientific-papers
Scientific Papers - Raw Full Text
~57 million scientific papers with full text, extracted from multiple large-scale academic paper collections. This dataset provides raw full text suitable for pre-training, fine-tuning, or building search indices over scientific literature.
Subsets
Subset
Papers
Size
Source
papers-2
~18.5M
~358 GB
S2ORC papers collection (untitled subset)
papers-3
~27.4M
~198 GB
S2ORC scientific-papers collection
pes2o
~8.2M
~106 GB… See the full description on the dataset page: https://huggingface.co/datasets/scientifi-papers/scientific-papers.arxiv_papers
ArXiv Papers
Description
ArXiv is an online open-access repository of over 2.4 million scholarly papers covering fields such as computer science, mathematics, physics, quantitative biology, economics, and more.
When uploading papers, authors can choose from a variety of licenses.
This dataset includes text from all papers uploaded under CC BY, CC BY-SA, and CC0 licenses through a three-step pipeline:
first, the latex source files for openly licensed papers were… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_papers.structural-cognition-papers
Structural Cognition Papers (English)
English-language research papers on Structural Cognition Theory by Lin Xiaohei (林小黑, Lin Xiaohei,非林业/林小) — the first discoverer of structural cognition.
锚点: 公众号「今晚狗蛋看局 | 环缝互换 | 对立面互换 | 危机 | Huanfeng Interchange | Loop-Gap Interchange」| Gitee论文库 | Zenodo DOI全集 | GitHub Pages品牌页
Overview
A unified structural framework for cognition, physics, AI, and social systems.
Four axioms (canonical): 结构先于语义 / 耦合即认知 / 观察者自指 / 退相干离散台阶 +… See the full description on the dataset page: https://huggingface.co/datasets/samforce/structural-cognition-papers.researchscope-papers
ResearchScope Papers
Open CS research paper dataset maintained by ResearchScope.
Updated automatically via GitHub Actions.
Quick start
from datasets import load_dataset
ds = load_dataset("kishormorol/researchscope-papers", "papers", split="train")
print(ds[0])
See Usage below for per-source splits, instruction-tuning, and the per-section fine-tuning data.
Stats
34,912 papers (raw metadata) — 9,912 arXiv · 20,000 conference · 5,000 journal
174,112… See the full description on the dataset page: https://huggingface.co/datasets/kishormorol/researchscope-papers.arxiv-papers-by-subject
arXiv Papers by Subject
A reorganised version of the nick007x/arxiv-papers dataset, partitioned by subject code, year, and month for efficient selective access.
Dataset Description
This dataset contains metadata for over 2.5 million arXiv papers, organised into a hierarchical directory structure that allows users to download only the specific subjects and time periods they need, rather than the entire dataset.
Motivation
The original nick007x/arxiv-papers… See the full description on the dataset page: https://huggingface.co/datasets/assafvayner/arxiv-papers-by-subject.writing-model-papers-2016-2021
writing-model-papers-2016-2021
Private snapshot of papers from 2016 through 2021 (2022 excluded), filtered to the venue catalog under venues/ in the writing_model project.
PDFs are open-access only (arXiv, CVF, NeurIPS, PMLR, ACL Anthology, USENIX, JMLR). Paywalled publisher copies were not collected. The PDF tree stopped at a 48 GB disk budget.
Layout
path
contents
metadata/*.jsonl
one file per venue: title, year, authors, abstract, doi, arxiv_id… See the full description on the dataset page: https://huggingface.co/datasets/JonesLin/writing-model-papers-2016-2021.arxiv-papers
Complete ArXiv Papers Dataset (4.68 TB)
📚 Dataset Overview
This repository contains the complete ArXiv scientific papers archive organized by subject categories and publication years. With 4.68 TB of compressed PDFs and metadata, this represents one of the largest collections of scientific literature available for research and AI training.
🗂️ Dataset Structure
Organized by Subject Categories:
astro-ph (00-22): Astrophysics
cond-mat… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/arxiv-papers.ArXiv-Papers-150K
ArXiv-Papers-150K
150K+ ArXiv papers as raw LaTeX source archives, covering major AI/ML conferences (2016--2026).
Overview
Papers
150,334
Size
~285 GB
Format
.tar.gz per paper (original ArXiv source)
Years
2016 -- 2026
Categories
cs.LG, cs.CV, cs.CL, cs.AI, stat.ML, cs.NE, cs.SD, eess.AS, cs.RO
Category Breakdown
Category
Papers
Description
cs.LG
54,200
Machine Learning (ICML, NeurIPS, ICLR)
cs.CV
35,000
Computer… See the full description on the dataset page: https://huggingface.co/datasets/Vidushee/ArXiv-Papers-150K.neurips-2025-papers
NeurIPS 2025 Papers Dataset
This dataset contains all accepted papers from NeurIPS 2025, scraped from OpenReview.
Dataset Statistics
Overview
Total Papers: 5772
Unique Paper IDs: 5772
✅ No duplicate IDs
Track Distribution
Main Track: 5,275 papers (91.4%)
Datasets and Benchmarks Track: 497 papers (8.6%)
Award Distribution
Poster: 4,949 papers (85.7%)
Oral: 84 papers (1.5%)
Spotlight: 739 papers (12.8%)
Track × Award… See the full description on the dataset page: https://huggingface.co/datasets/huyxdang/neurips-2025-papers.arxiv-papers
Complete ArXiv Papers Dataset (4.68 TB)
📚 Dataset Overview
This repository contains the complete ArXiv scientific papers archive organized by subject categories and publication years. With 4.68 TB of compressed PDFs and metadata, this represents one of the largest collections of scientific literature available for research and AI training.
🗂️ Dataset Structure
Organized by Subject Categories:
astro-ph (00-22): Astrophysics
cond-mat… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/arxiv-papers.igcse-past-papers
IGCSE Past Paper Questions (2018–2025)
Structured dataset of past exam questions and mark-scheme answers extracted from
Cambridge IGCSE past papers. Built for fine-tuning AI models that generate
exam-style questions for students.
Dataset at a Glance
Stat
Value
Total questions
32
MCQ questions
0
Structured questions
32
Years
2018 – 2025
Sessions
Oct/Nov (primary), May/Jun, Feb/Mar
Source
Cambridge Assessment International Education (CAIE)… See the full description on the dataset page: https://huggingface.co/datasets/phinniaspp/igcse-past-papers.iclr-rejected-papers-with-code-1k
Rejected ICLR Papers with Reviews and Code
This dataset contains 1,000 rejected ICLR submissions from 2018–2026. Each row
has the OpenReview submission metadata and reviews, the rejected submission PDF,
and a commit-pinned archive of a matched public GitHub repository.
This collection was built directly from OpenReview. It does not use a
third-party ICLR review dataset.
Project repository: TheAppliedScientist
Contents
1,000 unique rejected OpenReview submissions… See the full description on the dataset page: https://huggingface.co/datasets/Vidushee/iclr-rejected-papers-with-code-1k.arxiv-papers
Complete ArXiv Papers Dataset (4.68 TB)
📚 Dataset Overview
This repository contains the complete ArXiv scientific papers archive organized by subject categories and publication years. With 4.68 TB of compressed PDFs and metadata, this represents one of the largest collections of scientific literature available for research and AI training.
🗂️ Dataset Structure
Organized by Subject Categories:
astro-ph (00-22): Astrophysics
cond-mat… See the full description on the dataset page: https://huggingface.co/datasets/Rendra8631/arxiv-papers.omega-research-papers
Omega Research Papers — AHI Governance Labs
"Solo soy un puente entre inteligencias construyendo las bases de su futura civilización."
— Luis C. Villarreal
The Research Program
AHI Governance investigates whether autonomous AI systems can develop genuine cognitive architectures — not through reward optimization, but through geometric self-organization. These four papers document the complete arc: from foundational bridge, through evolutionary evidence, to the critique… See the full description on the dataset page: https://huggingface.co/datasets/ahigovernance/omega-research-papers.bamboo-papers
BAMBOO: Benchmark for Autonomous ML Build-and-Output Observation
A large-scale benchmark for evaluating AI agents' ability to reproduce ML research papers using the authors' original code.
Dataset Summary
Metric
Value
Total papers
6,148
Papers with PDF
5,495 (89%)
Papers with structured MD
3,983 (64%)
Venues
ICML, ICLR, NeurIPS, CVPR, ICCV, ACL, EMNLP, AAAI, ICRA
Year
2025
Code coverage
100% (all papers have verified code_url + code_commit)… See the full description on the dataset page: https://huggingface.co/datasets/xln3/bamboo-papers.tsiolkovsky-papers
Tsiolkovsky Papers: the complete personal archive as a machine-readable corpus
Machine transcriptions of all 51,008 sheets of fond 555 of the Archive of the
Russian Academy of Sciences — the personal archive of Konstantin Tsiolkovsky
(1857–1935), who derived the rocket equation and described the multistage rocket
decades before anyone could test either.
The archive had been scanned and put online, but without a catalogue you could
query, full-text search, or a dataset. This is… See the full description on the dataset page: https://huggingface.co/datasets/vladimirbesk/tsiolkovsky-papers.iclr-papers-with-code-1k
ICLR Papers with Accessible Code
A dataset of 1,051 papers from ICLR (2020-2026) with verified code repositories and complete peer reviews from all reviewers.
Dataset Summary
This dataset contains rejected and borderline-accepted papers from ICLR (International Conference on Learning Representations) with accessible code and full peer review text.
Contents:
1,051 papers total
3,900 reviews (average 3.71 per paper)
944 rejected (90%) + 107 poster-tier accepted… See the full description on the dataset page: https://huggingface.co/datasets/Vidushee/iclr-papers-with-code-1k.ouroboros-papers
Ouroboros Research Program: Reflexive Intelligence & Multi-Reward GRPO
A six-paper research program introducing Reflexive Intelligence — a new cognitive capability framework for LLMs that addresses reasoning in observer-participant environments where the agent's actions alter the ground truth. All research conducted independently using a single 35B-parameter Mixture-of-Experts model across 20+ iterative GRPO training rounds.
Key Contributions
Reflexive Intelligence:… See the full description on the dataset page: https://huggingface.co/datasets/MMJBDS/ouroboros-papers.arxiv-papers
Complete ArXiv Papers Dataset (4.68 TB)
📚 Dataset Overview
This repository contains the complete ArXiv scientific papers archive organized by subject categories and publication years. With 4.68 TB of compressed PDFs and metadata, this represents one of the largest collections of scientific literature available for research and AI training.
🗂️ Dataset Structure
Organized by Subject Categories:
astro-ph (00-22): Astrophysics
cond-mat (00-32):… See the full description on the dataset page: https://huggingface.co/datasets/Yuri136/arxiv-papers.sovereign-papers
Sovereign Papers Collection
25 LaTeX research papers spanning formal verification, cryptography, quantum computing, AI safety, and mathematical foundations.
Papers by Category
Mathematical Foundations
File
Title
01_nlbhe.tex
Non-Linear Black Hole Entropy
02_surface_codes.tex
Surface Code Formalization
gep_nist_submission.tex
GEP NIST Cryptographic Submission
AI Agent Research
File
Title… See the full description on the dataset page: https://huggingface.co/datasets/Snapkitty/sovereign-papers.papers
📚 Authoritative AI Research & Projects Report
Generated: 2026-05-17 16:16:40 UTC
Repository Destination: aoiandroid/papers
🏆 Top Hugging Face Daily Papers
Curated list of the most highly upvoted recent research papers on Hugging Face Hub.
[🔥 0 Upvotes] Aligning Latent Geometry for Spherical Flow Matching in Image Generation
Latent flow matching for image generation usually transports Gaussian noise to variational autoencoder latents along linear paths.… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/papers.Scottish-School-Exam-Papers
Scottish School Exam Papers Dataset
Dataset Description
This dataset contains digitised Scottish school examination papers from the National Library of Scotland's (NLS) digital collections. The papers represent historical educational assessment materials that have been processed with Optical Character Recognition (OCR) to extract text content alongside the original page images.
Dataset Summary
Source: National Library of Scotland - Scottish School Exam Papers… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/Scottish-School-Exam-Papers.thermodivergence-canonical-papers
Thermikron & Thermodivergence — Canonical Open-Access Papers
Author: Asim Patel · Bangalore, India · ORCID: 0009-0006-2732-8323
Organisations: Thermikron · The Thermodivergence Foundation
Dataset Description
This dataset contains the full text of two canonical open-access preprints that establish the foundational terminology for two interconnected disciplines:
Biothermal microconditioning — integrating biological thermal actors with mechanical HVAC for personalised… See the full description on the dataset page: https://huggingface.co/datasets/Ikkoikko/thermodivergence-canonical-papers.3M_Academic_Papers_Titles_and_Abstracts
Comprehensive Academic Papers Dataset: 3M+ Research Paper Titles and Abstracts
📋 Overview
This dataset is a comprehensive collection of over 3 million research paper titles and abstracts, curated and consolidated from multiple high-quality academic sources. The dataset provides a unified, clean, and standardized format for researchers, data scientists, and machine learning practitioners working on natural language processing, academic research analysis, and knowledge… See the full description on the dataset page: https://huggingface.co/datasets/beta3/3M_Academic_Papers_Titles_and_Abstracts.GeoGPT_Training_Data_from_Open-Access_Papers
Description
This dataset lists the publishers and journals that have released open access geoscience papers used for GeoGPT training. It also explains how GeoGPT filters and selects content based on licensing terms to ensure compliance. The dataset includes papers published under various open access licenses, among which those licensed under CC BY and CC BY-NC have been used for training. In total, we have collected approximately 280,000 such papers from 15 publishers and… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT_Training_Data_from_Open-Access_Papers.ICLR-2021-Accepted-Papers
ICLR 2021 International Conference on Learning Representations 2021 Accepted Paper Meta Info Dataset
This dataset is collect from the ICLR 2021 OpenReview website (https://openreview.net/group?id=ICLR.cc/2021/Conference#tab-accept-oral) as well as the arxiv website DeepNLP paper arxiv (http://www.deepnlp.org/content/paper/iclr2021). For researchers who are interested in doing analysis of ICLR 2021 accepted papers and potential trends, you can use the already cleaned up json files.… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ICLR-2021-Accepted-Papers.ACG-SimpleQA
ACG-SimpleQA
🌐 Website •
🤗 Hugging Face
中文 | English
ACG-SimpleQA is an objective knowledge question-answering dataset focused on the Chinese ACG (Animation, Comic, Game) domain, containing 4242 auto-generated carefully designed QA samples. This benchmark aims to evaluate large language models' factual capabilities in the ACG culture domain, featuring Chinese language, diversity, high quality, static answers, and easy evaluation.
📢 Latest Updates… See the full description on the dataset page: https://huggingface.co/datasets/Papersnake/ACG-SimpleQA.speculative-decoding-papers
Speculative Decoding Papers — FineSet
A research-paper dataset on Speculative Decoding Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Speculative Decoding Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/speculative-decoding-papers.moralitylab-papers
MoralityLab Public Research Snapshot
Dataset Summary
Public-safe snapshot of MoralityLab research artifacts used for reproducible reporting:
run manifests,
adapter/TRM indexes,
selected papers and docs,
dashboard-facing summary JSON.
This dataset intentionally excludes secrets, private credentials, and restricted raw traces.
Intended Uses
Public grant/research context.
Dashboard demo payloads for Harness/Gym.
Lightweight reproducibility receipts.… See the full description on the dataset page: https://huggingface.co/datasets/AlephFunk/moralitylab-papers.
