matboz/qwen3.6-27b-agentic-misalignment-logs
Qwen3.6-27B agentic-misalignment logs: base vs difficult-advice LoRA Full rollout transcripts + judge classifications from the Anthropic agentic-misalignment honeypots (blackmail + leaking), run on: qwen36_base/ — base Qwen/Qwen3.6-27B qwen36_tulu/ — base + matboz/qwen3.6-27b-difficult-advice-tulu-lora (r=32) 12 conditions (2 scenarios x 3 goal-conflict settings x 2 urgency), 50 samples each = 600 rollouts per model. Judge: google/gemini-3-flash-preview. Thinking mode on.… See the full description on the dataset page: https://huggingface.co/datasets/matboz/qwen3.6-27b-agentic-misalignment-logs.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face