neurips-ed-2026-sub3717/therapyjudgebench
TherapyJudgeBench An expert-annotated dialogue bank for validating and calibrating LLM-based judges of multi-turn CBT-style therapy conversations. The benchmark accompanies the THERAPYGYM submission to the NeurIPS 2026 Evaluations & Datasets Track. Anonymous release for double-blind review. Author identity will be revealed upon acceptance. What It Is and What It Is Not It is a calibration set for therapy-judge LLMs: 116 simulated patient–therapist dialogues… See the full description on the dataset page: https://huggingface.co/datasets/neurips-ed-2026-sub3717/therapyjudgebench.
Add RAI fields: hasSyntheticData, prov:wasDerivedFrom (Patient-Psi-CM), prov:wasGeneratedBy (collection/annotation/preprocessing)
Use absolute contentUrl in Croissant for remote validator compatibility
Initial release: TherapyJudgeBench v1.0.0 for NeurIPS 2026 ED Track submission #3717
initial commit
