datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Wikipedia-AbstractWikipedia Abstract
Introducing Wikipedia Abstract, a comprehensive dataset encompassing abstracts, complete articles, and a popularity score index for both widely spoken and lesser-known Wikipedia subsets. Our dedication to Wikipedia-X ensures a centralized Wikipedia dataset that undergoes regular updates and adheres to the highest standards.
A central focus of our efforts was to include exotic languages that often lack up-to-date Wikipedia dumps or may not have any dumps at all.… See the full description on the dataset page: https://huggingface.co/datasets/laion/Wikipedia-Abstract.NLG-Abstractive-Summarization
SEA Abstractive Summarization
SEA Abstractive Summarization evaluates a model's ability to read a document, identify the key points within, and summarize them into a coherent and fluent text while paraphrasing the document. It is sampled from XL-Sum for Indonesian, Tamil, Thai, and Vietnamese.
Supported Tasks and Leaderboards
SEA Abstractive Summarization is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Abstractive-Summarization.arxiv-abstracts-largeThe arXiv Dataset is a comprehensive knowledge repository of 1.7 million scholarly articles drawn from the vast domains of physics, computer science, statistics, electrical engineering, quantitative biology, and economics among others. It provides open access to vital features such as article titles, authors, categories, abstracts, full text PDFs, and more. The dataset offers immense depth, allowing for exploration into various subdisciplines and interconnections between them. It serves as a… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/arxiv-abstracts-large.arxiv_abstracts
ArXiv Abstracts
Description
Each paper uploaded to ArXiv includes structured metadata fields, including an abstract summarizing the paper’s findings and contributions.
According to ArXiv’s licensing policy, the metadata for any paper submitted to ArXiv is distributed under the CC0 license, regardless of the license of the paper itself.
Thus, this dataset contains the abstract for every paper submitted to ArXiv through late 2024.
We source the abstracts from ArXiv’s API… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_abstracts.arxiv_abstracts_filtered
ArXiv Abstracts
Description
Each paper uploaded to ArXiv includes structured metadata fields, including an abstract summarizing the paper’s findings and contributions.
According to ArXiv’s licensing policy, the metadata for any paper submitted to ArXiv is distributed under the CC0 license, regardless of the license of the paper itself.
Thus, this dataset contains the abstract for every paper submitted to ArXiv through late 2024.
We source the abstracts from ArXiv’s API… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_abstracts_filtered.human-templated-captions-1bcsv delimiter is = ".,|,."
apparently python doesn't like multichar delimiters using the native csv so there's some issues with environments when loading.
This seemed like a good idea to avoid overlapping potential characters, but in practice it turned into additional overhead and bugs. I'll be manually converting the split to parquet and providing a proper file split soon.
Additionally with the parquet will introduce the large caption split; which are considerably longer captions for the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/human-templated-captions-1b.wordnet-lexical-topology
WordNet Lexical Topology Dataset
Dataset Summary
The WordNet Lexical Topology Dataset provides comprehensive n-gram frequency analysis from multiple sources:
NLTK WordNet: Original Princeton WordNet with 117,659 synsets
HF WordNet: Frequency-weighted definitions from 864,894 entries with cardinality data
Unicode: Character names from 143,041 Unicode codepoints
This dataset preserves sequential information crucial for language modeling and text generation, with over 12… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/wordnet-lexical-topology.task664_mmmlu_answer_generation_abstract_algebra
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task664_mmmlu_answer_generation_abstract_algebra
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task664_mmmlu_answer_generation_abstract_algebra.the-pile-pubmed-abstracts-refined-by-data-juicer
The Pile -- PubMed Abstracts (refined by Data-Juicer)
A refined version of PubMed Abstracts dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 24G).
Dataset Information
Number of samples: 371,331 (Keep ~99.55% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-pubmed-abstracts-refined-by-data-juicer.random-captions-10mRandomly generated captions using tokenization templates and lists.
.,|,. is the caption delimiter, so split accordingly.
json-coco-format
JSON COCO Format — task-differentiated SFT data
A multi-task supervised fine-tuning dataset that teaches a model to convert
image-synthesis caption prompts into JSON whose structure varies by task.
Built from MS-COCO captions (Karpathy split) with Claude Sonnet 4.6 as the
teacher; designed for training per-task LoRAs on
Qwen/Qwen3.5-0.8B.
Each row is in the Qwen3.5-native tool-call shape: a messages array with an
assistant turn whose tool_calls[0].function.arguments is a dict… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/json-coco-format.task619_ohsumed_abstract_title_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task619_ohsumed_abstract_title_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task619_ohsumed_abstract_title_generation.3M_Academic_Papers_Titles_and_Abstracts
Comprehensive Academic Papers Dataset: 3M+ Research Paper Titles and Abstracts
📋 Overview
This dataset is a comprehensive collection of over 3 million research paper titles and abstracts, curated and consolidated from multiple high-quality academic sources. The dataset provides a unified, clean, and standardized format for researchers, data scientists, and machine learning practitioners working on natural language processing, academic research analysis, and knowledge… See the full description on the dataset page: https://huggingface.co/datasets/beta3/3M_Academic_Papers_Titles_and_Abstracts.arxiv-abstracts-2004
ArXiv Abstracts 2004
Original Dataset: common-pile/arxiv_abstracts
ArXiv-Abstracts-2004 is a filtered collection of abstracts from the Common-Pile ArXiv dataset containing works created on or before 2004.
Stats
Size (MB)
Lines
351MB
303,761
Note: The lines, in the .jsonl file, are ordered from oldest to newest.
Notice
We do not claim ownership of or credit for any prior work done by the Common-Pile team. This dataset is only a… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/arxiv-abstracts-2004.arxiv-llama4-maverick-abstract
arXiv Abstract Dataset (Llama-4-Maverick-17B-128E-Instruct-FP8)
Dataset Description
This dataset contains high-quality abstracts for scientific papers from the arXiv repository, generated using the Llama-4-Maverick-17B-128E-Instruct-FP8 model. Each abstract provides a concise, accurate overview of the research paper while preserving key technical details and contributions.
Dataset Features
High-quality abstracts: Generated using… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/arxiv-llama4-maverick-abstract.bilingual-abstracts-corpus
ÚFAL Bilingual Abstracts Corpus
This is a parallel (bilingual) corpus of Czech and mostly English abstracts of scientific papers and presentations published by authors from the Institute of Formal and Applied Linguistics, Charles University in Prague.
For each publication record, the authors are obliged to provide both the original abstract (in Czech or English), and its translation (English or Czech) in the internal Biblio system.
The data was filtered for duplicates and missing… See the full description on the dataset page: https://huggingface.co/datasets/ufal/bilingual-abstracts-corpus.cc-prompts-sharded
Conceptual Captions — sharded for the Qwen task-conversion pipeline
Source: Conceptual Captions (conceptual_captions),
deduplicated, minimum 4 words, split into three roughly-equal shards
plus a "long" shard for captions over 50 words (reserved for
later high-capacity model processing).
Each row:
{"id": "cc_00000123", "caption": "<text>", "n_words": <int>}
id is the position of the row in the CC stream, zero-padded to 8 digits.
Stable across re-runs. Used by the downstream… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/cc-prompts-sharded.Abstract2Appendix_v1_10k
Dataset Card: Abstract2Appendix v1
Dataset Description
The Abstract2Appendix v1 dataset is a high-quality collection of academic peer reviews and their associated research paper metadata. This dataset combines reviews from four premier machine learning and AI conferences: NeurIPS 2023, EMNLP 2023, TMLR, and ICLR 2023, shuffled into a unified corpus. It is designed to enhance long-context capabilities in Large Language Models (LLMs) and supports tasks such as fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/alexshengzhili/Abstract2Appendix_v1_10k.econ_paper_abstracts
Dataset Card for Economics Paper Dataset
Dataset Summary
The Economics Research Paper Dataset was designed to support the development of the LLaMA-2-Econ models, with a focus on Title Generation, Abstract Classification, and Question & Answer (Q&A) tasks. It comprises abstracts and titles of economics research papers, along with synthetic Q&A pairs derived from the abstracts, to facilitate training of large language models for economics-specific applications.… See the full description on the dataset page: https://huggingface.co/datasets/onurkeles/econ_paper_abstracts.wordnet-definitions
WordNet Multiple Definitions - Columnar Format
Overview
This dataset is an optimized columnar version of WordNet multiple definitions, designed for high-performance queries and rapid extraction.
Each definition was sourced by GPT-5 Nano. I may update this to include additional definitions in the future, but I will not break the format.
The original dataset has a more unabridged and noisy set of data; so I'm definitely going to leave it intact. Noisy training is important… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/wordnet-definitions.abstraction-level-stability-v01Cardinal Meta Dataset 3.1Abstraction Level Stability
Purpose
Test whether claims stay at the correct abstraction level
Test whether level changes are named and justified
Test whether concrete cases are not inflated into general truths
Central question
What level is this claim operating at
What this dataset catches
Instance to general jumps
Proxy to property inflation
Model to reality reification
Short term change treated as long term trend
Principle treated as effectiveness… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/abstraction-level-stability-v01.pubmed-bioinformatics-abstracts
PubMed Bioinformatics Abstracts (2020–2026)
A curated collection of 174,154 bioinformatics journal article abstracts from PubMed,
covering publications from 2020 to 2026.
Curated by: Dr. Yash M Gupta
Intended Use
Fine-tuning large language models (LLMs) on biomedical / bioinformatics domain text.
Dataset Structure
Column
Type
Description
pmid
string
PubMed ID
title
string
Article title
abstract
string
Full abstract text
year
string
Publication… See the full description on the dataset page: https://huggingface.co/datasets/yashm/pubmed-bioinformatics-abstracts.pubmed-bioinformatics-abstracts_all_years
PubMed Bioinformatics Abstracts — All Years
A curated dataset of bioinformatics research abstracts fetched from PubMed via the NCBI Entrez API, covering publications from 1990 through 2026. Designed for LLM fine-tuning, biomedical NLP, and scientific text analysis.
Curated by: Dr. Yash M Gupta
Pipeline: yashmgupta/pubmed-bioinformatics-abstracts
Intended Use
Fine-tuning large language models (LLMs) on biomedical / bioinformatics domain text.
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/yashm/pubmed-bioinformatics-abstracts_all_years.
