datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scientific-papersscientific-papers
Scientific Papers - Raw Full Text
~57 million scientific papers with full text, extracted from multiple large-scale academic paper collections. This dataset provides raw full text suitable for pre-training, fine-tuning, or building search indices over scientific literature.
Subsets
Subset
Papers
Size
Source
papers-2
~18.5M
~358 GB
S2ORC papers collection (untitled subset)
papers-3
~27.4M
~198 GB
S2ORC scientific-papers collection
pes2o
~8.2M
~106 GB… See the full description on the dataset page: https://huggingface.co/datasets/scientifi-papers/scientific-papers.scientific_papersScientific papers datasets contains two sets of long and structured documents.
The datasets are obtained from ArXiv and PubMed OpenAccess repositories.
Both "arxiv" and "pubmed" have two features:
- article: the body of the document, pagragraphs seperated by "/n".
- abstract: the abstract of the document, pagragraphs seperated by "/n".
- section_names: titles of sections, seperated by "/n".LimAgents_limitation_data_scientific_papers_with_cited_papers
LimAgents Data
This dataset contains scientific paper metadata and extracted limitation information prepared for use with LLM Agents.The data comes from NeurIPS 2021–2022 papers and related OpenReview reviews, enriched with Cited in and Cited by information.
Dataset Structure
The repository contains two main directories:
1. NeurIPS_21_22_Lim_OPR_with_cited_in_by_papers
This directory includes one JSON file per paper. Each file contains:
title: Original paper… See the full description on the dataset page: https://huggingface.co/datasets/iaadlab/LimAgents_limitation_data_scientific_papers_with_cited_papers.scientific_papers_arxiv_promptsourcescientific_papers_pubmed_promptsourcescientific_papers_dummyScientific papers datasets contains two sets of long and structured documents.
The datasets are obtained from ArXiv and PubMed OpenAccess repositories.
Both "arxiv" and "pubmed" have two features:
- article: the body of the document, pagragraphs seperated by "/n".
- abstract: the abstract of the document, pagragraphs seperated by "/n".
- section_names: titles of sections, seperated by "/n".scientific_papers-archive
Dataset Card for "scientific_papers"
Dataset Summary
Scientific papers datasets contains two sets of long and structured documents.
The datasets are obtained from ArXiv and PubMed OpenAccess repositories.
Both "arxiv" and "pubmed" have two features:
article: the body of the document, paragraphs separated by "/n".
abstract: the abstract of the document, paragraphs separated by "/n".
section_names: titles of sections, separated by "/n".
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/scillm/scientific_papers-archive.scientific-papersscientific_papers_DANCERprocessed_scientific_papers
Preprocessing used
Removing Stopwords, Removing Punctuation
Data Fields
The data fields are the same among all splits.
pubmed
article: a string feature.
abstract: a string feature.
section_names: a string feature.
Data Splits
name
train
validation
test
pubmed
119924
6633
6658
scientific_papers_citation_scores
Dataset Summary
This dataset comprises an array of scientific papers, each paper is associated with a series of scores.
These scores quantify the number of citations each paper has received.
The data regarding the papers and their citations were sourced from OpenCitations, a comprehensive and accessible online database of scholarly citations (available at https://opencitations.net/).
How are these scores calculated?
Imagine a tree where papers are nodes and citations… See the full description on the dataset page: https://huggingface.co/datasets/JoaoCoelho/scientific_papers_citation_scores.scientific_papersscientific-papers-dataset
Scientific Papers Dataset
Scientific papers, whitepapers and documentation.
Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente.
cross-domain-scientific-papers-tripletsprocessed_pubmed_scientific_papers
Dataset Card for "processed_pubmed_scientific_papers"
More Information needed
autoeval-eval-scientific_papers-pubmed-c3b6df-51381145313
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: Samuel-Fipps/t5-efficient-large-nl36_fine_tune_sum_V2
Dataset: scientific_papers
Config: pubmed
Split: train
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @NessTechIntl for evaluating this model.
scientific_papers-cleaned
Dataset Card for Scientific Papers - Cleaned
This dataset card provides information about the Scientific Papers - Cleaned dataset, a refined and curated version of scientific papers for research and natural language processing tasks.
Dataset Details
Dataset Description
The Scientific Papers - Cleaned dataset is a cleaned version of a larger collection of scientific papers. It is designed to provide structured input-output pairs for tasks such as text… See the full description on the dataset page: https://huggingface.co/datasets/GilbertKrantz/scientific_papers-cleaned.scientific_papers_esscientific_papers_enscientific-papers-3.5-withpromptarmanc_scientific_papers_pubmed_1024Summarize-Scientific-Papers-Processed
Scientific summarization dataset
Original dataset: https://huggingface.co/datasets/inference-net/Small-Summarized-Paper-Dataset
The current version of dataset is a part of the original dataset, processed into a readable and usable json format with relevant data of each paper.
scientific_papersscientific_papers_en_escollabllm-multiturn-scientific-papers-summarizationscientific_papers_esscientific_papersScientific papers datasets contains two sets of long and structured documents.
The datasets are obtained from ArXiv and PubMed OpenAccess repositories.
Both "arxiv" and "pubmed" have two features:
- article: the body of the document, pagragraphs seperated by "/n".
- abstract: the abstract of the document, pagragraphs seperated by "/n".
- section_names: titles of sections, seperated by "/n".ScientificPapersscientific_papers_summary_dataset
