datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Wikipedia-AbstractWikipedia Abstract
Introducing Wikipedia Abstract, a comprehensive dataset encompassing abstracts, complete articles, and a popularity score index for both widely spoken and lesser-known Wikipedia subsets. Our dedication to Wikipedia-X ensures a centralized Wikipedia dataset that undergoes regular updates and adheres to the highest standards.
A central focus of our efforts was to include exotic languages that often lack up-to-date Wikipedia dumps or may not have any dumps at all.… See the full description on the dataset page: https://huggingface.co/datasets/laion/Wikipedia-Abstract.hotpotqa-fr-abstracts
HotpotQA-fr (abstracts)
65 565 questions multi-sauts en français, au format de
HotpotQA (Yang et al., 2018). Les questions sont
construites directement sur Wikipédia français, sans traduction. L'annotation
humaine du jeu original est remplacée par une génération par modèle de langue
suivie d'une validation automatique par ablations.
Chaque contexte contient dix paragraphes : les deux abstracts nécessaires à la
réponse et huit distracteurs. Tous les paragraphes sont des abstracts… See the full description on the dataset page: https://huggingface.co/datasets/Mvanypersele/hotpotqa-fr-abstracts.biomedical-abstract-qa-training-pool
Biomedical abstract question answering training pool
Public research questions over biomedical abstracts from three datasets, read at the pinned
revisions named below and laid out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 395882 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file
question
the research question, as its source publishes it… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/biomedical-abstract-qa-training-pool.Pes2o-Abstract-XIntroducing Pes2o-X, also known as Pes2o-Abstract-X, a derived dataset from the original Pes2o dataset released by Allen AI. The Pes2o dataset aimed to provide a large corpus of open-access research papers, including both abstracts and full text. However, it required pre-processing before the abstracts could be used for training or fine-tuning machine learning models.
At LAION AI, we initiated a project called X, focusing on developing high-quality training corpora from scratch, reorganising… See the full description on the dataset page: https://huggingface.co/datasets/laion/Pes2o-Abstract-X.econ_paper_abstracts
Dataset Card for Economics Paper Dataset
Dataset Summary
The Economics Research Paper Dataset was designed to support the development of the LLaMA-2-Econ models, with a focus on Title Generation, Abstract Classification, and Question & Answer (Q&A) tasks. It comprises abstracts and titles of economics research papers, along with synthetic Q&A pairs derived from the abstracts, to facilitate training of large language models for economics-specific applications.… See the full description on the dataset page: https://huggingface.co/datasets/onurkeles/econ_paper_abstracts.wordnet-definitions
WordNet Multiple Definitions - Columnar Format
Overview
This dataset is an optimized columnar version of WordNet multiple definitions, designed for high-performance queries and rapid extraction.
Each definition was sourced by GPT-5 Nano. I may update this to include additional definitions in the future, but I will not break the format.
The original dataset has a more unabridged and noisy set of data; so I'm definitely going to leave it intact. Noisy training is important… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/wordnet-definitions.qa-from-abstract-grapheneabstractive-qa-llm-eval
Abstractive QA for LLM evaluation
A Hebrew-language dataset of 132 question-answer records, built for testing
whether a language model's generated answer is grounded in its source
document rather than hallucinated. Each record has a question, a
source_text, and a reference_answer whose individual claims are each
backed by exact quoted spans from source_text. Use it to check your own
model the same way: generate an answer, break it into claims, and see
whether each claim traces… See the full description on the dataset page: https://huggingface.co/datasets/HebArabNlpProject/abstractive-qa-llm-eval.
