CoolFace
Datasetpublicgated

HelpMum-Personal/MamaBench

Dataset Card for MamaBench Dataset Summary MamaBench is a counterfactual clinical benchmark for evaluating the robustness of large language models on maternal and child health diagnostic reasoning. It consists of 217 counterfactual case pairs, each pairing an original clinical vignette with a systematically perturbed counterfactual variant, designed to test whether a model's diagnostic reasoning is sensitive to clinically meaningful changes rather than relying on… See the full description on the dataset page: https://huggingface.co/datasets/HelpMum-Personal/MamaBench.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
2likes36downloads

No commit history came back for main. The revision may not exist, or the source declined the request.