datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MedThinkVQA
MedThinkVQA
MedThinkVQA is an expert-annotated benchmark for multi-image diagnostic reasoning in radiology. Unlike prior medical VQA benchmarks that typically contain at most one image per case, MedThinkVQA requires models to extract evidence from each image, integrate cross-view information, and perform differential-diagnosis reasoning.
Links
GitHub: https://github.com/benluwang/MedThinkVQA
Leaderboard: https://benluwang.github.io/MedThinkVQA/
Submission Guide:… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/MedThinkVQA.bionlp2004[BioNLP2004 NER dataset](https://aclanthology.org/W04-1213.pdf)bionlp_st_2013_cgthe Cancer Genetics (CG) is a event extraction task and a main task of the BioNLP Shared Task (ST) 2013.
The CG task is an information extraction task targeting the recognition of events in text,
represented as structured n-ary associations of given physical entities. In addition to
addressing the cancer domain, the CG task is differentiated from previous event extraction
tasks in the BioNLP ST series in addressing a wide range of pathological processes and multiple
levels of biological organization, ranging from the molecular through the cellular and organ
levels up to whole organisms. Final test set submissions were accepted from six teamsNoteAid-READMEThe Datasets contains the jargon terms, lay definitions, general definitions for different stages in our REAME pipeline. To comply with fair use of law~\footnote{\url{https://www.copyright.gov/fair-use/}}, We used GPT-3.5 to paraphrase the lay definitions as shown in synthetic_data_creation.ipynb. We used GPT-4o-mini to paraphrase the EHRs as shown in synthetic_EHR_creation.ipynb. We asked the LLM(gpt-4o-mimi) to edit the original sentence but make sure to keep the main terms unchanged. Check… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/NoteAid-README.bioinstruct
Dataset Card for BioInstruct
GitHub repo: https://github.com/bio-nlp/BioInstruct
Dataset Summary
BioInstruct is a dataset of 25k instructions and demonstrations generated by OpenAI's GPT-4 engine in July 2023.
This instruction data can be used to conduct instruction-tuning for language models (e.g. Llama) and make the language model follow biomedical instruction better.
Improvements of Llama on 9 common BioMedical tasks are shown in the result section.
Taking… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/bioinstruct.bionlp_st_2011_geThe BioNLP-ST GE task has been promoting development of fine-grained information extraction (IE) from biomedical
documents, since 2009. Particularly, it has focused on the domain of NFkB as a model domain of Biomedical IE.
The GENIA task aims at extracting events occurring upon genes or gene products, which are typed as "Protein"
without differentiating genes from gene products. Other types of physical entities, e.g. cells, cell components,
are not differentiated from each other, and their type is given as "Entity".MedQA-MM
MedQA-MM Identifier Release
Paper repository ·
Hugging Face dataset
MedQA-MM is a 1,000-item shortcut-mitigated medical multimodal multiple-choice benchmark constructed from MedThinkVQA, MedXpertQA-MM, and the Health and Medicine portion of MMMU. This public release is intentionally identifier-only.
It does not contain source questions, answer choices, gold answers, images, clinical text, or repaired payloads. It provides stable source locators, pinned source revisions, and a… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/MedQA-MM.bionlp_st_2013_pcthe Pathway Curation (PC) task is a main event extraction task of the BioNLP shared task (ST) 2013.
The PC task concerns the automatic extraction of biomolecular reactions from text.
The task setting, representation and semantics are defined with respect to pathway
model standards and ontologies (SBML, BioPAX, SBO) and documents selected by relevance
to specific model reactions. Two BioNLP ST 2013 participants successfully completed
the PC task. The highest achieved F-score, 52.8%, indicates that event extraction is
a promising approach to supporting pathway curation efforts.NoteAid_Chatbotbionlp_st_2011_relThe Entity Relations (REL) task is a supporting task of the BioNLP Shared Task 2011.
The task concerns the extraction of two types of part-of relations between a
gene/protein and an associated entity.bionlp_st_2013_groGRO Task: Populating the Gene Regulation Ontology with events and
relations. A data set from the bio NLP shared tasks competition from 2013MedQA-CS-ExamBenchmarking LLMs Clinical Skills for Patient-Centered Diagnostics and Documentation
Project github: https://github.com/bio-nlp/MedQA-CS
MedQA-CS-Student dataset: https://huggingface.co/datasets/bio-nlp-umass/MedQA-CS-Student
BioNLPbionlp_st_2019_bbThe task focuses on the extraction of the locations and phenotypes of
microorganisms from PubMed abstracts and full-text excerpts, and the
characterization of these entities with respect to reference knowledge
sources (NCBI taxonomy, OntoBiotope ontology). The task is motivated by
the importance of the knowledge on biodiversity for fundamental research
and applications in microbiology.bionlp_shared_task_2009The BioNLP Shared Task 2009 was organized by GENIA Project and its corpora were curated based
on the annotations of the publicly available GENIA Event corpus and an unreleased (blind) section
of the GENIA Event corpus annotations, used for evaluation.combined_bionlp_task_dataset_model_cardsbionlp_st_2011_epiThe dataset of the Epigenetics and Post-translational Modifications (EPI) task
of BioNLP Shared Task 2011.Synth-SBDH
Dataset Card for Synth-SBDH
Synth-SBDH is a collection of 8,767 synthetic examples with annotations for 15 SBDH categories. SBDH annotations include information such as presence, period and annotation rationale.
Dataset Description
Synth-SBDH is a novel synthetic SBDH dataset that mimics EHR notes.
Repository: Codes to reproduce experiments
Paper: Link
Point of Contact: Avijit Mitra
Dataset Structure
Data Instances
Some examples from… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/Synth-SBDH.eval-gliner2-ner-bionlp2004-boundary-smoothing-validationeval-gliner2-ner-bionlp2004-boundary-smoothing-testMedQA-CS-StudentTask_models_BioNLP_results_GPTSynth-PPD
Dataset Card for Synth-PPD
Synth-PPD is a collection of 7,579 synthetic examples annotated for parole, probation or unclear (Label). Annotations also include information such as presence, period and annotation reasoning.
Dataset Structure
Data Instances
Some examples from synth_data_mc_mtl_train.json looks as follows.
{
'Text': 'He has been facing challenges in complying with the terms of his probation, including community service.',
'idx': 263… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/Synth-PPD.BioNLP2021MEDIQA @ NAACL-BioNLP 2021 -- Task 2: Multi-answer summarization
https://sites.google.com/view/mediqa2021
Biomedical Summarization Data
The MEDIQA-AnS Dataset could be used for training.bionlp2004revise_bigbio_task_models_w_BioNLPbionlp2[BioNLP2004 NER dataset](https://aclanthology.org/W04-1213.pdf)BioNLP_Filtered
Dataset Card for BioNLP2004 Filtered
This dataset is an adaptation of the original tner/bionlp2004 dataset, specifically tailored for our specific project use-case which focuses on CellLine and CellType entities.
Dataset Description
The original BioNLP2004 dataset is a named entity recognition (NER) dataset in the biomedical domain, annotated with various entity types such as DNA, Protein, Cell_type, Cell_line, and RNA.
This adapted version, OTAR3088/BioNLP_Filtered, has… See the full description on the dataset page: https://huggingface.co/datasets/OTAR3088/BioNLP_Filtered.bionlp_st_2011_idThe dataset of the Infectious Diseases (ID) task of
BioNLP Shared Task 2011.bionlp_st_2013_geThe BioNLP-ST GE task has been promoting development of fine-grained
information extraction (IE) from biomedical
documents, since 2009. Particularly, it has focused on the domain of
NFkB as a model domain of Biomedical IE
