scientific papers
scientific-papersscientific-papers
Scientific Papers - Raw Full Text
~57 million scientific papers with full text, extracted from multiple large-scale academic paper collections. This dataset provides raw full text suitable for pre-training, fine-tuning, or building search indices over scientific literature.
Subsets
Subset
Papers
Size
Source
papers-2
~18.5M
~358 GB
S2ORC papers collection (untitled subset)
papers-3
~27.4M
~198 GB
S2ORC scientific-papers collection
pes2o
~8.2M
~106 GB… See the full description on the dataset page: https://huggingface.co/datasets/scientifi-papers/scientific-papers.scientific_papersScientific papers datasets contains two sets of long and structured documents.
The datasets are obtained from ArXiv and PubMed OpenAccess repositories.
Both "arxiv" and "pubmed" have two features:
- article: the body of the document, pagragraphs seperated by "/n".
- abstract: the abstract of the document, pagragraphs seperated by "/n".
- section_names: titles of sections, seperated by "/n".LimAgents_limitation_data_scientific_papers_with_cited_papers
LimAgents Data
This dataset contains scientific paper metadata and extracted limitation information prepared for use with LLM Agents.The data comes from NeurIPS 2021–2022 papers and related OpenReview reviews, enriched with Cited in and Cited by information.
Dataset Structure
The repository contains two main directories:
1. NeurIPS_21_22_Lim_OPR_with_cited_in_by_papers
This directory includes one JSON file per paper. Each file contains:
title: Original paper… See the full description on the dataset page: https://huggingface.co/datasets/iaadlab/LimAgents_limitation_data_scientific_papers_with_cited_papers.scientific_papers_arxiv_promptsourcescientific_papers_pubmed_promptsource
