datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LRM-Safety-evaluation-parsedDiNa-LRM-SD35m-HPSv3-Preprocess-Datalrm-safety-eval
Chain of Risk — LRM Safety Evaluation Dataset
⚠️ Content Warning: This dataset contains potentially harmful, unsafe, or unethical prompts
and model responses collected strictly for safety research purposes.
Dataset Summary
This dataset accompanies the paper:
Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering
Xiaomin Li, Jianheng Hou, Zheyuan Deng, Zhiwei Zhang, Taoran Li, Binghang Lu, Bing Hu, Yunhan Zhao… See the full description on the dataset page: https://huggingface.co/datasets/HJH2CMD/lrm-safety-eval.ifeval-lrm
IFEval Dataset Card
Dataset Description
IFEval is an instruction-following evaluation benchmark consisting of verifiable natural language instructions. Each example specifies one or more constraints that a model must satisfy in its output (e.g., include/exclude phrases, follow a format, respect length or style constraints). In this project, IFEval is used to evaluate both:
instruction following in the reasoning trace (RT), and
instruction following in the final… See the full description on the dataset page: https://huggingface.co/datasets/haritzpuerto/ifeval-lrm.lrm_safety_alignment_sftlrm_safety_alignment_dpolrm_logsLRMsafety
