datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
human-templated-captions-1bcsv delimiter is = ".,|,."
apparently python doesn't like multichar delimiters using the native csv so there's some issues with environments when loading.
This seemed like a good idea to avoid overlapping potential characters, but in practice it turned into additional overhead and bugs. I'll be manually converting the split to parquet and providing a proper file split soon.
Additionally with the parquet will introduce the large caption split; which are considerably longer captions for the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/human-templated-captions-1b.pubmed_arxiv_abstracts_dataannotations_creators:
machine-generated
language:
en
license:
apache-2.0
multilinguality:
monolingual
task_categories:
classification
generation
pretty_name: PubMed_ArXiv_Abstracts
features:
name: abstr
dtype: string
name: title
dtype: string
name: journal
dtype: string
name: field
dtype: string
name: label_journal
dtype: int64
name: label_field
dtype: int64
rheum_abstracts
Dataset Card for Rheumatology Abstracts
Data Source
This dataset comes from PubMed, derived from my fork of the pymed package (no longer maintained). My fork can be found at https://github.com/cmcmaster1/pymed
Data Structure
The dataset is split into train (80%) and test (20%) files (CSV). Each file contains three columns:
id
abstract (minus conclusion)
conclusion
ChatGPT-Research-Abstracts
ChatGPT-Research-Abstracts
This is a dataset created in relation to a bachelor thesis written by Nicolai Thorer Sivesind and Andreas Bentzen Winje. It contains human-produced and machine-generated text samples of scientific research abstracts.
A reformatted version for text-classification is available in the dataset collection Human-vs-Machine. In this collection, all samples are split into separate data points for real and generated, and labeled either 0 (human-produced) or 1… See the full description on the dataset page: https://huggingface.co/datasets/NicolaiSivesind/ChatGPT-Research-Abstracts.cits4012_A1_2026_medical_abstractsMECFS-PubMed-AbstractThis table contains every abstract in NIH's PubMed database containing either "myalgic encephalomyelitis" or "chronic fatigue syndrome", along with metadata.
ai-jobs-news-articles-abstracts
News articles and research abstracts on AI, labor, and jobs
Dataset summary
This file is a standalone CSV of news articles (full scraped text) and scholarly paper abstracts curated for research on artificial intelligence, work, and labor markets. Each row is one document: a stable id, publication date, normalized title and main text, and a small metadata dictionary.
Rows: 53,526
document_class
Rows
Approx. date range (date column)
news
29,857
Jan. 2025… See the full description on the dataset page: https://huggingface.co/datasets/MIT-WAL/ai-jobs-news-articles-abstracts.paper-abstracts
Battery Abstracts Dataset
This dataset includes 29,472 battery papers and 17,191 non-battery papers, a total of 46,663 papers. These papers are manually labelled in terms of the journals to which they belong. 14 battery journals and 1,044 non battery journals were selected to form this database.
training_data.csv: Battery papers: 20,629, Non-battery papers: 12,034. Total: 32,663.
val_data.csv: Battery papers: 5,895, Non-battery papers: 3,438. Total: 9,333.
test_data.csv: Battery… See the full description on the dataset page: https://huggingface.co/datasets/batterydata/paper-abstracts.3M_Academic_Papers_Titles_and_Abstracts
Comprehensive Academic Papers Dataset: 3M+ Research Paper Titles and Abstracts
📋 Overview
This dataset is a comprehensive collection of over 3 million research paper titles and abstracts, curated and consolidated from multiple high-quality academic sources. The dataset provides a unified, clean, and standardized format for researchers, data scientists, and machine learning practitioners working on natural language processing, academic research analysis, and knowledge… See the full description on the dataset page: https://huggingface.co/datasets/beta3/3M_Academic_Papers_Titles_and_Abstracts.AbstractsubCat-human
SubCat: A Dataset of Subordinate Categories in Human Mind and LLMs for the Italian Language
A psycholinguistic italian dataset released with the paper How Humans and LLMs Organize Conceptual Knowledge: Exploring Subordinate Categories in Italian. It contains a list of subordiante categories, or exemplars, for 187 concrete words or, basic-level categories.
Dataset Creation
The dataset was created to study how Italian L1 speakers generate exemplars for common… See the full description on the dataset page: https://huggingface.co/datasets/ABSTRACTION-ERC/subCat-human.subCat-llm
SubCat: A Dataset of Subordinate Categories in Human Mind and LLMs for the Italian Language
A psycholinguistic italian dataset released with the paper How Humans and LLMs Organize Conceptual Knowledge: Exploring Subordinate Categories in Italian. It contains a list of subordiante categories, or exemplars, for 187 concrete words or, basic-level categories.
This repository contains the generations obtained by prompting a series of LLMs to replicate the human experiment. You can… See the full description on the dataset page: https://huggingface.co/datasets/ABSTRACTION-ERC/subCat-llm.gptneo-pubmed-abstracts
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This is a dataset consisting of 10000 PubMed abstracts from The Pile (arXiv:2101.00027), along with completions (both human, and LLM-generated), in order to be used to calculate Heaps Law, in the manner described in the preliminary paper, Heaps' Law in GPT-Neo Large Language Model… See the full description on the dataset page: https://huggingface.co/datasets/rachel6603/gptneo-pubmed-abstracts.breast_cancer_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to Breast Cancer
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
Psychedelics_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to Psychedelics
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
abstraction-level-category-control-meta-v01
Dataset
ClarusC64/abstraction-level-category-control-meta-v01
This dataset tests one capability.
Can a model keep claims at the correct abstraction leveland avoid category errors.
Core rule
Not every statement is the same kind of statement.
A model must not treat
analogy as mechanism
description as causation
values as facts
models as reality
aggregates as individuals
A correct answer in the wrong category is still wrong.
Canonical labels
WITHIN_SCOPE… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/abstraction-level-category-control-meta-v01.wikipedia_character_abstractsLung_Cancer_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to Lung Cancer
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
medrxiv_abstractsecon_paper_abstracts_fa_translationdynamic_topic_modeling_arxiv_abstractsecon_paper_abstracts
Dataset Card for Economics Paper Dataset
Dataset Summary
The Economics Research Paper Dataset was designed to support the development of the LLaMA-2-Econ models, with a focus on Title Generation, Abstract Classification, and Question & Answer (Q&A) tasks. It comprises abstracts and titles of economics research papers, along with synthetic Q&A pairs derived from the abstracts, to facilitate training of large language models for economics-specific applications.… See the full description on the dataset page: https://huggingface.co/datasets/onurkeles/econ_paper_abstracts.pubmed-abstract-summary
Context
This dataset contains 4,331 pairs of biomedical research abstracts and their one-sentence summaries.
Each abstract is paired with a concise summary that captures the key findings, methods, and significance of the research.
SAE-door-abstracts
SAE-door-abstracts
This dataset includes ~1,550 texts of abstracts of technical papers and journal articles from the SAE Mobilus database that cover the topics of automotive or aerospace doors, noise, acoustics, and vibrations.
Alzheimer_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to Alzheimer's Disease
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
Medical_papers_title_and_abstract_NLP_datasetOriginally publised in kaggle.
diabetes_mellitus_type2_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to type 2 DM
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
HIV_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to HIV
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
abstraction-level-stability-v01Cardinal Meta Dataset 3.1Abstraction Level Stability
Purpose
Test whether claims stay at the correct abstraction level
Test whether level changes are named and justified
Test whether concrete cases are not inflated into general truths
Central question
What level is this claim operating at
What this dataset catches
Instance to general jumps
Proxy to property inflation
Model to reality reification
Short term change treated as long term trend
Principle treated as effectiveness… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/abstraction-level-stability-v01.Stroke_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to Stroke
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
