datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pubmed_arxiv_abstracts_dataannotations_creators:
machine-generated
language:
en
license:
apache-2.0
multilinguality:
monolingual
task_categories:
classification
generation
pretty_name: PubMed_ArXiv_Abstracts
features:
name: abstr
dtype: string
name: title
dtype: string
name: journal
dtype: string
name: field
dtype: string
name: label_journal
dtype: int64
name: label_field
dtype: int64
rheum_abstracts
Dataset Card for Rheumatology Abstracts
Data Source
This dataset comes from PubMed, derived from my fork of the pymed package (no longer maintained). My fork can be found at https://github.com/cmcmaster1/pymed
Data Structure
The dataset is split into train (80%) and test (20%) files (CSV). Each file contains three columns:
id
abstract (minus conclusion)
conclusion
ChatGPT-Research-Abstracts
ChatGPT-Research-Abstracts
This is a dataset created in relation to a bachelor thesis written by Nicolai Thorer Sivesind and Andreas Bentzen Winje. It contains human-produced and machine-generated text samples of scientific research abstracts.
A reformatted version for text-classification is available in the dataset collection Human-vs-Machine. In this collection, all samples are split into separate data points for real and generated, and labeled either 0 (human-produced) or 1… See the full description on the dataset page: https://huggingface.co/datasets/NicolaiSivesind/ChatGPT-Research-Abstracts.cits4012_A1_2026_medical_abstracts3M_Academic_Papers_Titles_and_Abstracts
Comprehensive Academic Papers Dataset: 3M+ Research Paper Titles and Abstracts
📋 Overview
This dataset is a comprehensive collection of over 3 million research paper titles and abstracts, curated and consolidated from multiple high-quality academic sources. The dataset provides a unified, clean, and standardized format for researchers, data scientists, and machine learning practitioners working on natural language processing, academic research analysis, and knowledge… See the full description on the dataset page: https://huggingface.co/datasets/beta3/3M_Academic_Papers_Titles_and_Abstracts.ai-jobs-news-articles-abstracts
News articles and research abstracts on AI, labor, and jobs
Dataset summary
This file is a standalone CSV of news articles (full scraped text) and scholarly paper abstracts curated for research on artificial intelligence, work, and labor markets. Each row is one document: a stable id, publication date, normalized title and main text, and a small metadata dictionary.
Rows: 53,526
document_class
Rows
Approx. date range (date column)
news
29,857
Jan. 2025… See the full description on the dataset page: https://huggingface.co/datasets/MIT-WAL/ai-jobs-news-articles-abstracts.paper-abstracts
Battery Abstracts Dataset
This dataset includes 29,472 battery papers and 17,191 non-battery papers, a total of 46,663 papers. These papers are manually labelled in terms of the journals to which they belong. 14 battery journals and 1,044 non battery journals were selected to form this database.
training_data.csv: Battery papers: 20,629, Non-battery papers: 12,034. Total: 32,663.
val_data.csv: Battery papers: 5,895, Non-battery papers: 3,438. Total: 9,333.
test_data.csv: Battery… See the full description on the dataset page: https://huggingface.co/datasets/batterydata/paper-abstracts.wikipedia_character_abstractsPsychedelics_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to Psychedelics
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
gptneo-pubmed-abstracts
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This is a dataset consisting of 10000 PubMed abstracts from The Pile (arXiv:2101.00027), along with completions (both human, and LLM-generated), in order to be used to calculate Heaps Law, in the manner described in the preliminary paper, Heaps' Law in GPT-Neo Large Language Model… See the full description on the dataset page: https://huggingface.co/datasets/rachel6603/gptneo-pubmed-abstracts.breast_cancer_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to Breast Cancer
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
Lung_Cancer_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to Lung Cancer
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
medrxiv_abstractsecon_paper_abstracts
Dataset Card for Economics Paper Dataset
Dataset Summary
The Economics Research Paper Dataset was designed to support the development of the LLaMA-2-Econ models, with a focus on Title Generation, Abstract Classification, and Question & Answer (Q&A) tasks. It comprises abstracts and titles of economics research papers, along with synthetic Q&A pairs derived from the abstracts, to facilitate training of large language models for economics-specific applications.… See the full description on the dataset page: https://huggingface.co/datasets/onurkeles/econ_paper_abstracts.dynamic_topic_modeling_arxiv_abstractsecon_paper_abstracts_fa_translationSAE-door-abstracts
SAE-door-abstracts
This dataset includes ~1,550 texts of abstracts of technical papers and journal articles from the SAE Mobilus database that cover the topics of automotive or aerospace doors, noise, acoustics, and vibrations.
Alzheimer_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to Alzheimer's Disease
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
diabetes_mellitus_type2_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to type 2 DM
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
HIV_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to HIV
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
Stroke_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to Stroke
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
Type_2_Diabetes_Mellitus_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to type 2 DM
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
abstractsBrain_Tumor_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to Brain Tumors
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
midas-abstracts
Dataset Card for MIDAS Abstracts Dataset
This dataset contains abstracts from papers associated with MIDAS members. Current coverage is for papers from 2015-2024 and obtained from the MIDAS website.
acl-2024-long-abstractsarxiv-abstracts-methodsMTL-abstractsTHESES-abstractsAcute_Lymphoblastic_Leukemia_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to acute lymphoblastic leukemia
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
