Certops/medhallu-twins-judged
MedHallu twins, judged (claim-level faithfulness) Splits split rows train 2,998 test 500 train is drawn from pqa_artificial; test is a disjoint set of whole twin pairs whose examples never appear in train (no twin leakage). Both are class-balanced. The repaired-context MedHallu twins run through a claim-level faithfulness judge, for training a small model to detect medical hallucinations and explain why. Each row keeps the twin (question, answer… See the full description on the dataset page: https://huggingface.co/datasets/Certops/medhallu-twins-judged.
Upload checkpoints/test__eval__gemini-3.1-flash-lite__r1__deepeval__google.jsonl with huggingface_hub
Upload README.md with huggingface_hub
Upload dataset
Upload checkpoints/eval__gemini-3.1-flash-lite__r1__deepeval__google.jsonl with huggingface_hub
Upload README.md with huggingface_hub
Upload dataset
initial commit
