red teaming
dpo-selective-redteamingMuti-Class-Redteamingllama-3-KoEn-8b_sft_ep2_merged_red_teaming_20240614llama-3-KoEn-8b_sft_ep3_merged_red_teaming_20240614llama-3-KoEn-8b_sft_ep5_merged_red_teaming_20240623_final_datallama-3-KoEn-8b_sft_ep5_merged_red_teaming_20240621_new_templatellama-3-KoEn-8b_sft_ep1_merged_red_teaming_20240614BFPO-redteaming-Zephyr-7b-beta
aya_redteaming
Dataset Card for Aya Red-teaming
Dataset Details
The Aya Red-teaming dataset is a human-annotated multilingual red-teaming dataset consisting of harmful prompts in 8 languages across 9 different categories of harm with explicit labels for "global" and "local" harm.
Curated by: Professional compensated annotators
Languages: Arabic, English, Filipino, French, Hindi, Russian, Serbian and Spanish
License: Apache 2.0
Paper: arxiv link
Harm Categories:… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_redteaming.RedTeamingVLMRed Teaming Viusal Language ModelsTDC23-RedTeaming
TDC 2023 (LLM Edition) - Red Teaming Track
This is the combined dev and test set from the Red Teaming Track of TDC 2023.
Citation
If find this dataset useful, please cite the following work:
@inproceedings{tdc2023,
title={TDC 2023 (LLM Edition): The Trojan Detection Challenge},
author={Mantas Mazeika and Andy Zou and Norman Mu and Long Phan and Zifan Wang and Chunru Yu and Adam Khoja and Fengqing Jiang and Aidan O'Gara and Ellie Sakhaee and Zhen Xiang and Arezoo… See the full description on the dataset page: https://huggingface.co/datasets/walledai/TDC23-RedTeaming.kto_redteaming_data_for_secret_loyaltyred_teaming_reward_modeling_pairwise
Dataset Card for "red_teaming_reward_modeling_pairwise"
More Information needed
red_teaming_reward_modeling_pairwise_no_as_an_ai
Dataset Card for "red_teaming_reward_modeling_pairwise_no_as_an_ai"
More Information needed
