nlm
Datasets
All datasets matching “nlm”MedFact-SynthTo train Med-V1, we construct MedFact-Synth, a large-scale synthetic training set including 1.5 million instances.
Each instance contains: a synthetic claim to be verified, a source article serving as evidence, a rationale explaining the verification, and a 5-point Likert-scale verdict, ranging from strong contradiction (-2) and partial contradiction (-1) to neutral (0), partial agreement (+1), and strong agreement (+2).
To build this dataset, we begin by sampling one million articles from… See the full description on the dataset page: https://huggingface.co/datasets/nlm-dir/MedFact-Synth.nlm-cxr_reportsMedFact-Bench
[!Note]
The inclusion of datasets does not imply endorsement or agreement to their content by the authors or their employers. The datasets were selected based on prior work in the field of claim verification.
To evaluate Med-V1, we curate MedFact-Bench, a benchmark comprising five biomedical verification datasets: SciFact, HealthVer, MedAESQA, PubMedQA-Fact (re-purposed PubMedQA), and BioASQ-Fact (re-purposed BioASQ).
Across all datasets, each instance consists of a claim–source pair… See the full description on the dataset page: https://huggingface.co/datasets/nlm-dir/MedFact-Bench.NLMCXR
Dataset Card for "NLMCXR"
More Information needed
nlm_geneNLM-Gene consists of 550 PubMed articles, from 156 journals, and contains more than 15 thousand unique gene names, corresponding to more than five thousand gene identifiers (NCBI Gene taxonomy). This corpus contains gene annotation data from 28 organisms. The annotated articles contain on average 29 gene names, and 10 gene identifiers per article. These characteristics demonstrate that this article set is an important benchmark dataset to test the accuracy of gene recognition algorithms both on multi-species and ambiguous data. The NLM-Gene corpus will be invaluable for advancing text-mining techniques for gene identification tasks in biomedical text.nlmchem
Dataset Details
Dataset Description
NLM-Chem is a new resource for chemical entity recognition in PubMed full text literature.
Curated by:
License: CC BY 4.0
Dataset Sources
data source
publication
Citation
BibTeX:
@article{Islamaj2021,
author = {Islamaj, R. and Leaman, R. and Kim, S. and Lu, Z.},
title = {NLM-Chem, a new resource for chemical entity recognition in PubMed full text literature},
journal = {Nature Scientific Data},
volume = {8}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nlmchem.
