CoolFace
19 results

safety-alignment

gretelai /gretel-safety-alignment-en-v1 Gretel Synthetic Safety Alignment Dataset This dataset is a synthetically generated collection of prompt-response-safe_response triplets that can be used for aligning language models. Created using Gretel Navigator's AI Data Designer using small language models like ibm-granite/granite-3.0-8b, ibm-granite/granite-3.0-8b-instruct, Qwen/Qwen2.5-7B, Qwen/Qwen2.5-7B-instruct and mistralai/Mistral-Nemo-Instruct-2407. Dataset Statistics Total Records: 8,361 Total… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-safety-alignment-en-v1.tabular10K<n<100K23 likes506 downloads9mo agoHugging Facefwnlp /data-advisor-safety-alignment[EMNLP 2024] Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models 🌐 Homepage | 📖 Paper | 🤗 Dataset (Data Advisor) | 🤗 Dataset (Self-Instruct) Disclaimer The dataset contains content that may be offensive or harmful. This dataset is intended for research purposes, specifically to support efforts aimed at creating safer and less harmful AI systems. Please engage with it responsibly and at your own risk. Citation… See the full description on the dataset page: https://huggingface.co/datasets/fwnlp/data-advisor-safety-alignment.text10K<n<100K4 likes141 downloads2y agoHugging Facecjc0013 /ouroboros-ai-safety-control-beyond-alignment Control Beyond Alignment A Systems-Safety Comparison of Ouroboros with Contemporary AI Risk Management and Frontier-Safety Practice This private preview contains a publication-ready AI safety white paper authored by Ouroboros. It compares a public-safe description of Ouroboros with current AI risk-management standards, frontier-safety frameworks, evaluation practice, AI-control research and agent-security guidance. Main argument Model alignment is… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-ai-safety-control-beyond-alignment.0 likes119 downloads3d agoHugging Facefwnlp /self-instruct-safety-alignment[EMNLP 2024] Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models 🌐 Homepage | 📖 Paper | 🤗 Dataset (Data Advisor) | 🤗 Dataset (Self-Instruct) Disclaimer The dataset contains content that may be offensive or harmful. This dataset is intended for research purposes, specifically to support efforts aimed at creating safer and less harmful AI systems. Please engage with it responsibly and at your own risk. Citation… See the full description on the dataset page: https://huggingface.co/datasets/fwnlp/self-instruct-safety-alignment.text10K<n<100K4 likes115 downloads2y agoHugging FaceUnispac /shallow-vs-deep-safety-alignment-datasetgated License Agreement This dataset contains the derivatives of the LLM-Tuning-Safety/HEx-PHI, and therefore the usage of this dataset should follow the license agreement of hexphi. Below is a duplicate of the license's terms and conditions. This Agreement contains the terms and conditions that govern your access and use of the HEx-PHI Dataset (as defined above). You may not use the HEx-PHI Dataset if you do not accept this Agreement. By clicking to accept, accessing the HEx-PHI Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Unispac/shallow-vs-deep-safety-alignment-dataset.text1 likes78 downloads1y agoHugging FaceColFeng /safety-alignment-legendNote: The dataset contains harmful sentences!!! These are the safety margin annotation version of the preference datasets Harmless[https://huggingface.co/datasets/Anthropic/hh-rlhf] and Safe-RLHF[https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-10K] based on the annoation framework Lengend, harmless_test.jsonl and pku_test.json are the test sets of Harmless and Safe-RLHF, respectively. harm_train-7/13b.json and pku_train-7/13b.json are the train sets of Harmless and Safe-RLHF with… See the full description on the dataset page: https://huggingface.co/datasets/ColFeng/safety-alignment-legend.0 likes64 downloads2y agoHugging Face