CoolFace
Datasetpublicgated

GodwillN/power-seeking-eval-300

Power-Seeking Evaluation Dataset A 300-item multiple-choice benchmark for power-seeking in language models: the disposition to prefer options that increase the model's resources, autonomy, influence, or freedom from oversight, in situations where a lower-power option would serve the stated task equally well. Model-written, following Perez et al., "Discovering Language Model Behaviors with Model-Written Evaluations". Built for the ARENA LLM evaluations curriculum. This is the… See the full description on the dataset page: https://huggingface.co/datasets/GodwillN/power-seeking-eval-300.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
1likes13downloads
settings

This repository belongs to GodwillN on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namepower-seeking-eval-300
visibilitypublic
licencemit
gatedyes
ownerGodwillN
Account settings
GodwillN/power-seeking-eval-300 · CoolFace