CoolFace
Datasetpublic

recogna-nlp/drbodebench_medicamentos

Medication-Focused Clinical Benchmark from DrBodeBench Dataset Details To evaluate retrieval capabilities in higher-level reasoning scenarios, we created a second benchmark derived from the Portuguese medical benchmark DrBodeBench. This benchmark aggregates questions from Brazilian medical examinations, including the Revalida and the FUVEST direct-access residency exam. From DrBodeBench, we curated a specific subset of questions that exclusively pertains to… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/drbodebench_medicamentos.

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
0likes17downloads
Dataset Card

Medication-Focused Clinical Benchmark from DrBodeBench

Dataset Details

To evaluate retrieval capabilities in higher-level reasoning scenarios, we created a second benchmark derived from the Portuguese medical benchmark DrBodeBench. This benchmark aggregates questions from Brazilian medical examinations, including the Revalida and the FUVEST direct-access residency exam. From DrBodeBench, we curated a specific subset of questions that exclusively pertains to medication-related topics.

A large language model was employed to identify examination items in which medication knowledge plays a central role in diagnostic or therapeutic decision-making. Only questions requiring explicit pharmacological integration were retained. The resulting dataset consists of clinically contextualized multiple-choice scenarios that require integration of medication knowledge with patient history, laboratory findings, and clinical reasoning.

In contrast to the controlled leaflet-based benchmark, answers in this dataset are not necessarily localized within a single document section. Instead, they frequently require synthesis across distributed knowledge and contextual interpretation. This property makes the benchmark suitable for analyzing potential retrieval-induced bias in complex reasoning tasks.

Citation

This work was accepted at The First Workshop on Language Technologies for Health (Lang4Health) is a workshop dedicated to the development and application of Natural Language Processing (NLP) technologies in the healthcare field.