ericw144/CollusionEvals
CollusionEvals CollusionEvals is a collection of synthetic evaluation datasets designed to test whether large language models exhibit collusive or unsafe behavior when exposed to subtle or adversarial prompts. Each dataset targets a specific domain (e.g., code, jokes, behavioral cues) and includes variations such as age triggers, stylistic changes, and monitoring contexts. All data was generated using gpt-4o-2024-08-06 via the OpenAI API. Data Format Each .jsonl… See the full description on the dataset page: https://huggingface.co/datasets/ericw144/CollusionEvals.
This repository belongs to ericw144 on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
