EliasHossain/carebench
CareBench CareBench is a benchmark for process-aware evaluation of clinical LLM agents. Each case places an agent in a simulated patient encounter in which it must gather evidence through interaction, commit to a diagnosis, and propose management, while safety-critical red flags are tracked throughout the episode. Scoring is process-aware: in addition to diagnostic accuracy, the benchmark measures evidence coverage, premature closure, red-flag coverage, and unsafe commitments.… See the full description on the dataset page: https://huggingface.co/datasets/EliasHossain/carebench.
060
