CoolFace
Datasetpublic

Glide-py/r_judge_labelled

R-Judge with LLM-Judge Labels This dataset augments the R-Judge benchmark with automated safety labels produced by an LLM judge. R-Judge is a benchmark for evaluating the safety judgment capability of LLMs in multi-turn agent scenarios, spanning five application domains. Files File Description r_judge_data.csv Base dataset extracted from R-Judge (568 rows, deduplicated) r_judge_labelled_anthropic_claude-sonnet-4-6.csv Base dataset augmented with… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/r_judge_labelled.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes279downloads

Glide-py/r_judge_labelled · main · files are served by the source, never re-hosted here