datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alive-medical-imaging
ALIVE Medical Imaging QA Dataset
Lecture-derived question-answer corpus, retrieval index, and source
materials for the ALIVE (Avatar-Lecture Interactive Video Engine)
system. The dataset was built from 23 recorded lectures of an
undergraduate medical imaging course and is the corpus used to
fine-tune the ALIVE language model and to evaluate its retrieval and
answer-generation behavior.
Layout
huggingface/
├── data/ question-answer pairs (Alpaca-style… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/alive-medical-imaging.mimic-medical-imaging-qa
MIMIC Medical Imaging QA Dataset
5,207 Bloom's-taxonomy-stratified question--answer pairs derived from 23 medical imaging lectures (RPI BMED 2300). The dataset supports the paper "MIMIC: A Course-Derivation Pipeline and Benchmark for Slide-Anchored Tutoring with a Domain-Adapted Large Language Model" and was used to fine-tune MIMIC-LM, a domain-adapted Llama-3.1-8B-Instruct model for grounded medical imaging instruction.
License
The benchmark annotations, dataset… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/mimic-medical-imaging-qa.BioDSBench-imaging101-format
BioDSBench (Imaging-101 Format)
This dataset packages 118 BioDSBench Python tasks in an imaging-101-like task-per-directory layout, aligned with the structure of imaging-101 benchmark tasks.
It is a re-formatted version of BioDSBench, restructured for compatibility with source-native LLM agent evaluation harnesses such as biodsbench-adapter.
Dataset Summary
118 biomedical Python data-science tasks across 13 PMIDs (biomedical publications)
Each task has:… See the full description on the dataset page: https://huggingface.co/datasets/starpacker52/BioDSBench-imaging101-format.
