abstracts
BunnyBosz_-_llama-3.2-3b-fine-tuned-model-Mut-effect-pred-v3-abstracts-only-ggufBunnyBosz_-_llama-3.2-3b-fine-tuned-model-Mut-effect-pred-v1-abstracts-only-ggufBunnyBosz_-_llama-3.2-3b-fine-tuned-model-Mut-effect-pred-v4-abstracts-only-ggufBunnyBosz_-_llama-3.2-3b-fine-tuned-model-Mut-effect-pred-v2-abstracts-only-ggufabstract-sim-sentenceabstract-sim-queryabstract-sim-sentence-pubmedabstract-sim-query-pubmed
arxiv-abstracts-2021
Dataset Card for arxiv-abstracts-2021
Dataset Summary
A dataset of metadata including title and abstract for all arXiv articles up to the end of 2021 (~2 million papers).
Possible applications include trend analysis, paper recommender engines, category prediction, knowledge graph construction and semantic search interfaces.
In contrast to arxiv_dataset, this dataset doesn't include papers submitted to arXiv after 2021 and it doesn't require any external download.… See the full description on the dataset page: https://huggingface.co/datasets/gfissore/arxiv-abstracts-2021.pubmed-abstracts-36M
OmniBioAI PubMed Abstracts — 37.8M+
The most comprehensive open collection of
PubMed biomedical abstracts.
Stats
37,846,388 abstracts (full PubMed coverage)
150 biomedical domains
56 general corpus chunks
207 total files
JSONL.gz format (human readable)
FREE and open access
Coverage
Complete PubMed database as of 2026.
Format
Each line = one abstract in JSON:
{"pmid": "...", "title": "...",
"abstract": "...", "authors": [...]… See the full description on the dataset page: https://huggingface.co/datasets/omnibioai/pubmed-abstracts-36M.arxiv-abstracts-largeThe arXiv Dataset is a comprehensive knowledge repository of 1.7 million scholarly articles drawn from the vast domains of physics, computer science, statistics, electrical engineering, quantitative biology, and economics among others. It provides open access to vital features such as article titles, authors, categories, abstracts, full text PDFs, and more. The dataset offers immense depth, allowing for exploration into various subdisciplines and interconnections between them. It serves as a… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/arxiv-abstracts-large.pubtator3_abstracts
PubTator3 dataset
PubTator3 annotations.
The dataset contains titles, abstracts, publication year, and annotation data for each annotation predicted by PubTator3.
If an abstract was split into multiple parts in the PubTator3 archive files, they have been joined so each publication has exactly one abstract.
In addition to PubTator3 data, this has been enriched with reference data pulled from the PubMed XML files.
Update
This dataset has been updated November 18th… See the full description on the dataset page: https://huggingface.co/datasets/dconnell/pubtator3_abstracts.arxiv_abstracts
ArXiv Abstracts
Description
Each paper uploaded to ArXiv includes structured metadata fields, including an abstract summarizing the paper’s findings and contributions.
According to ArXiv’s licensing policy, the metadata for any paper submitted to ArXiv is distributed under the CC0 license, regardless of the license of the paper itself.
Thus, this dataset contains the abstract for every paper submitted to ArXiv through late 2024.
We source the abstracts from ArXiv’s API… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_abstracts.papers-with-abstracts
[!CAUTION]
This dataset will not be updated. It corresponds to the last available public snapshot of the data, retrieved on July 29th, 2025.
