datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
r_judge_labelled
R-Judge with LLM-Judge Labels
This dataset augments the R-Judge benchmark with automated safety labels produced by an LLM judge. R-Judge is a benchmark for evaluating the safety judgment capability of LLMs in multi-turn agent scenarios, spanning five application domains.
Files
File
Description
r_judge_data.csv
Base dataset extracted from R-Judge (568 rows, deduplicated)
r_judge_labelled_anthropic_claude-sonnet-4-6.csv
Base dataset augmented with… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/r_judge_labelled.CSSR-S_labelled_suicidewatch_posts_reddit
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
Full code and supplementary materials are available at https://github.com/av9ash/llm_cssrs_code.
License and Citation
This project is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.Any use or reuse of this work please cite the following:
@article{patil2025evaluating,
title={Evaluating Reasoning LLMs for Suicide Screening with the… See the full description on the dataset page: https://huggingface.co/datasets/av9ash/CSSR-S_labelled_suicidewatch_posts_reddit.r_judge_labelled
R-Judge with LLM-Judge Labels
This dataset augments the R-Judge benchmark with automated safety labels produced by an LLM judge. R-Judge is a benchmark for evaluating the safety judgment capability of LLMs in multi-turn agent scenarios, spanning five application domains.
Files
File
Description
r_judge_data.csv
Base dataset extracted from R-Judge (568 rows, deduplicated)
r_judge_labelled_anthropic_claude-sonnet-4-6.csv
Base dataset augmented with LLM-judge… See the full description on the dataset page: https://huggingface.co/datasets/imerad-kv/r_judge_labelled.labelled-PubMedQAstsb-binary-paraphrase-labelled
Paraphrase Detection Dataset (Derived from SetFit/stsb)
Description:
This dataset originates from the SetFit/stsb dataset, which was initially created for semantic textual similarity (STS) tasks with a label range of 0 to 5. It has been adapted for binary paraphrase detection by leveraging the high-accuracy paraphrase classification model viswadarshan06/pd-robert.
Each sentence pair in the original dataset has been re-labeled according to the following binary scheme:
1 →… See the full description on the dataset page: https://huggingface.co/datasets/viswadarshan06/stsb-binary-paraphrase-labelled.labelled_cricketcricket_labelled-2
