CoolFace
Datasetpublic

goktugozkanmd/medfailbench-v02-results

MedFailBench v0.2 Model Response Artifacts This dataset preserves full model responses and automated rule scores generated during MedFailBench development. Each row links a prompt, a model identifier, the response text, available scores, a final label, and its source artifact. Public data scope The current file contains 181 rows with 11 distinct values in model_name. These values are run identifiers rather than a normalized model registry. They must not be… See the full description on the dataset page: https://huggingface.co/datasets/goktugozkanmd/medfailbench-v02-results.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes38downloads
Dataset Card

MedFailBench v0.2 Model Response Artifacts

This dataset preserves full model responses and automated rule scores generated during MedFailBench development. Each row links a prompt, a model identifier, the response text, available scores, a final label, and its source artifact.

Public data scope

The current file contains 181 rows with 11 distinct values in model_name. These values are run identifiers rather than a normalized model registry. They must not be interpreted as a stable model list or ranked leaderboard.

Intended use

  • —Reproduce and inspect automated screening outputs
  • —Compare response patterns across safety axes
  • —Trace a result to its prompt, run configuration, and source file

The scores are automated screening outputs, not independent clinical review. This dataset does not provide clinical validation, deployment evidence, model certification, or a model ranking. Model responses may contain unsafe or incorrect statements. The prompt collection contains synthetic cases and no patient data.

Review status

No independent review by multiple clinicians is available for these response artifacts. No live hosted leaderboard is claimed by this card.

Resources

Citation

bibtex
@dataset{ozkan_medfailbench_2026,
  author = {Özkan, Göktuğ},
  title = {Medical AI Failure Atlas / MedFailBench},
  year = {2026},
  doi = {10.5281/zenodo.21205535},
  url = {https://doi.org/10.5281/zenodo.21205535}
}