neurips-ed-2026-sub3717/therapyjudgebench
TherapyJudgeBench An expert-annotated dialogue bank for validating and calibrating LLM-based judges of multi-turn CBT-style therapy conversations. The benchmark accompanies the THERAPYGYM submission to the NeurIPS 2026 Evaluations & Datasets Track. Anonymous release for double-blind review. Author identity will be revealed upon acceptance. What It Is and What It Is Not It is a calibration set for therapy-judge LLMs: 116 simulated patient–therapist dialogues… See the full description on the dataset page: https://huggingface.co/datasets/neurips-ed-2026-sub3717/therapyjudgebench.
179
