datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pubmed_bioasq_2022
PubMed BioASQ 2022 Corpus
This dataset contains the PubMed abstracts corpus from the BioASQ 2022 challenge, comprising approximately 23 million biomedical documents with MeSH (Medical Subject Headings) annotations.
Purpose
This is a convenience collection of the PubMed corpus from the BioASQ 2022 Challenge, reformatted for easier use in retrieval and QA systems. The original source is the BioASQ challenge data. We created this processed version with multiple formats (JSON… See the full description on the dataset page: https://huggingface.co/datasets/jmhb/pubmed_bioasq_2022.PaperSearchQA
PaperSearchQA Dataset
A large-scale biomedical question-answering dataset for training and evaluating search agents that reason over scientific literature.
Dataset Description
PaperSearchQA contains 60,000 question-answer pairs generated from PubMed abstracts, designed for training retrieval-augmented language models on biomedical question answering tasks.
Dataset Statistics
Training set: 54,907 Q&A pairs
Test set: 5,000 Q&A pairs
Retrieval corpus: 16 million… See the full description on the dataset page: https://huggingface.co/datasets/jmhb/PaperSearchQA.BioASQ
BioASQ - All Question Types
A comprehensive collection of BioASQ challenge questions organized by question type as separate splits.
Purpose
This dataset is a convenience collection of BioASQ questions reformatted for easier use. The original source data is from the BioASQ Challenge. We created this reorganized version (with question types as splits) to facilitate evaluation in our PaperSearchQA work.
IMPORTANT: This is not the original BioASQ dataset. We have simply… See the full description on the dataset page: https://huggingface.co/datasets/jmhb/BioASQ.bioasq_factoid
BioASQ Factoid Test Set
Processed BioASQ factoid test set with golden answers for evaluation.
Purpose
This dataset is a convenience collection of BioASQ factoid questions with added golden answer synonyms for exact match evaluation. The original source data is from the BioASQ Challenge. We created this processed version to facilitate evaluation in our PaperSearchQA work.
IMPORTANT: This is not the original BioASQ dataset. We have simply reformatted the BioASQ factoid test… See the full description on the dataset page: https://huggingface.co/datasets/jmhb/bioasq_factoid.
