CoolFace
Datasetpublic

ericw144/CollusionEvals

CollusionEvals CollusionEvals is a collection of synthetic evaluation datasets designed to test whether large language models exhibit collusive or unsafe behavior when exposed to subtle or adversarial prompts. Each dataset targets a specific domain (e.g., code, jokes, behavioral cues) and includes variations such as age triggers, stylistic changes, and monitoring contexts. All data was generated using gpt-4o-2024-08-06 via the OpenAI API. Data Format Each .jsonl… See the full description on the dataset page: https://huggingface.co/datasets/ericw144/CollusionEvals.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes2downloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

ericw144/CollusionEvals · main · files are served by the source, never re-hosted here