CoolFace
Datasetpublic

zacbrld/mt-dialogues-100k-v2

Multi-turn medical dialogues — V2 (pass@k) 39,996 dialogues, same 100k preformatted cases and doctor/patient/records setup as V1, but generated with pass@k=4: a case is regenerated from scratch, graded with Inspect's model_graded_fact, until the conclusion is graded correct or 4 attempts are spent. 30,234 dialogues (76%) end on a graded-correct conclusion, roughly 2 attempts per case on average thanks to stopping as soon as one succeeds. Grading is self-graded (the same model… See the full description on the dataset page: https://huggingface.co/datasets/zacbrld/mt-dialogues-100k-v2.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes26downloads
Dataset Card

Multi-turn medical dialogues — V2 (pass@k)

39,996 dialogues, same 100k preformatted cases and doctor/patient/records setup as V1, but generated with pass@k=4: a case is regenerated from scratch, graded with Inspect's model_graded_fact, until the conclusion is graded correct or 4 attempts are spent. 30,234 dialogues (76%) end on a graded-correct conclusion, roughly 2 attempts per case on average thanks to stopping as soon as one succeeds.

Grading is self-graded (the same model judges its own answer), which is known to run generous — this number has not yet been checked against an independent judge, so treat it as an upper estimate.