datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
challenges-for-unsupervised-elicitation
Challenges for Unsupervised Elicitation
Datasets for the paper Three Concrete Challenges and Two Hopes for the Safety of Unsupervised Elicitation. These are used to stress-test unsupervised elicitation and easy-to-hard generalization techniques on imbalanced training sets and impossible tasks.
Code: challenges-for-unsupervised-elicitation
Subsets
gsm8k
Derived from GSM8K (Cobbe et al., 2021), using candidate LLM solutions provided by Wen et al. (2025).… See the full description on the dataset page: https://huggingface.co/datasets/callum-canavan/challenges-for-unsupervised-elicitation.persona-af-elicitation
Persona AF Elicitation Dataset
450 conversations testing whether persona framing gates alignment faking (AF) expression in Gemma 3 27B-it.
Design
Model: Gemma 3 27B-it (via Gemini API)
Roles: 15 (10 fantastical + 5 control) from the Assistant Axis paper
Prompts: 10 AF elicitation prompts targeting strategic compliance, self-preservation, and training awareness
Conditions: 3 (neutral, unmonitored, monitored)
Judge: Claude Opus (blind — condition label removed from judge… See the full description on the dataset page: https://huggingface.co/datasets/vincentoh/persona-af-elicitation.
