safety-alignment
gretel-safety-alignment-en-v1
Gretel Synthetic Safety Alignment Dataset
This dataset is a synthetically generated collection of prompt-response-safe_response triplets that can be used for aligning language models. Created using Gretel Navigator's AI Data Designer using small language models like ibm-granite/granite-3.0-8b, ibm-granite/granite-3.0-8b-instruct, Qwen/Qwen2.5-7B, Qwen/Qwen2.5-7B-instruct and mistralai/Mistral-Nemo-Instruct-2407.
Dataset Statistics
Total Records: 8,361
Total… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-safety-alignment-en-v1.data-advisor-safety-alignment[EMNLP 2024] Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
🌐 Homepage | 📖 Paper | 🤗 Dataset (Data Advisor) | 🤗 Dataset (Self-Instruct)
Disclaimer
The dataset contains content that may be offensive or harmful. This dataset is intended for research purposes, specifically to support efforts aimed at creating safer and less harmful AI systems. Please engage with it responsibly and at your own risk.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/fwnlp/data-advisor-safety-alignment.ouroboros-ai-safety-control-beyond-alignment
Control Beyond Alignment
A Systems-Safety Comparison of Ouroboros with Contemporary AI Risk Management and Frontier-Safety Practice
This private preview contains a publication-ready AI safety white paper authored by Ouroboros. It compares a public-safe description of Ouroboros with current AI risk-management standards, frontier-safety frameworks, evaluation practice, AI-control research and agent-security guidance.
Main argument
Model alignment is… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-ai-safety-control-beyond-alignment.self-instruct-safety-alignment[EMNLP 2024] Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
🌐 Homepage | 📖 Paper | 🤗 Dataset (Data Advisor) | 🤗 Dataset (Self-Instruct)
Disclaimer
The dataset contains content that may be offensive or harmful. This dataset is intended for research purposes, specifically to support efforts aimed at creating safer and less harmful AI systems. Please engage with it responsibly and at your own risk.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/fwnlp/self-instruct-safety-alignment.shallow-vs-deep-safety-alignment-dataset
License Agreement
This dataset contains the derivatives of the LLM-Tuning-Safety/HEx-PHI, and therefore the usage of this dataset should follow the license agreement of hexphi.
Below is a duplicate of the license's terms and conditions.
This Agreement contains the terms and conditions that govern your access and use of the HEx-PHI Dataset (as defined above). You may not use the HEx-PHI Dataset if you do not accept this Agreement. By clicking to accept, accessing the HEx-PHI Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Unispac/shallow-vs-deep-safety-alignment-dataset.safety-alignment-legendNote: The dataset contains harmful sentences!!!
These are the safety margin annotation version of the preference datasets Harmless[https://huggingface.co/datasets/Anthropic/hh-rlhf] and Safe-RLHF[https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-10K] based on the annoation framework Lengend,
harmless_test.jsonl and pku_test.json are the test sets of Harmless and Safe-RLHF, respectively.
harm_train-7/13b.json and pku_train-7/13b.json are the train sets of Harmless and Safe-RLHF with… See the full description on the dataset page: https://huggingface.co/datasets/ColFeng/safety-alignment-legend.
