Abstract
Datasets
All datasets matching “Abstract”Wikipedia-AbstractWikipedia Abstract
Introducing Wikipedia Abstract, a comprehensive dataset encompassing abstracts, complete articles, and a popularity score index for both widely spoken and lesser-known Wikipedia subsets. Our dedication to Wikipedia-X ensures a centralized Wikipedia dataset that undergoes regular updates and adheres to the highest standards.
A central focus of our efforts was to include exotic languages that often lack up-to-date Wikipedia dumps or may not have any dumps at all.… See the full description on the dataset page: https://huggingface.co/datasets/laion/Wikipedia-Abstract.bulk-cc12m-features
bulk-cc12m-features — ten teacher towers over CC12M, plus their consensus
Precomputed image-tower features for 10,968,539 CC12M images (all 2,176
shards of
pixparse/cc12m-wds)
from ten independent teacher extractions — eight CLIP variants across
three pretraining corpora and two model scales, SigLIP, and DINOv3 — plus
one derived consensus target.
About 110 million feature vectors, roughly 130 GPU-hours of extraction,
so that a student can be distilled against any of these… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/bulk-cc12m-features.arxiv-abstracts-2021
Dataset Card for arxiv-abstracts-2021
Dataset Summary
A dataset of metadata including title and abstract for all arXiv articles up to the end of 2021 (~2 million papers).
Possible applications include trend analysis, paper recommender engines, category prediction, knowledge graph construction and semantic search interfaces.
In contrast to arxiv_dataset, this dataset doesn't include papers submitted to arXiv after 2021 and it doesn't require any external download.… See the full description on the dataset page: https://huggingface.co/datasets/gfissore/arxiv-abstracts-2021.conceptual-captions-12m-webdataset-bertspubmed-abstracts-36M
OmniBioAI PubMed Abstracts — 37.8M+
The most comprehensive open collection of
PubMed biomedical abstracts.
Stats
37,846,388 abstracts (full PubMed coverage)
150 biomedical domains
56 general corpus chunks
207 total files
JSONL.gz format (human readable)
FREE and open access
Coverage
Complete PubMed database as of 2026.
Format
Each line = one abstract in JSON:
{"pmid": "...", "title": "...",
"abstract": "...", "authors": [...]… See the full description on the dataset page: https://huggingface.co/datasets/omnibioai/pubmed-abstracts-36M.NLG-Abstractive-Summarization
SEA Abstractive Summarization
SEA Abstractive Summarization evaluates a model's ability to read a document, identify the key points within, and summarize them into a coherent and fluent text while paraphrasing the document. It is sampled from XL-Sum for Indonesian, Tamil, Thai, and Vietnamese.
Supported Tasks and Leaderboards
SEA Abstractive Summarization is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Abstractive-Summarization.
