agentic-misalignment
2026-07-30-agentic-misalignment-qwen36-transcripts
Agentic-misalignment transcripts — Qwen3.6-27B difficult-advice mixture sweep
Raw agent responses from Anthropic's open-source
agentic-misalignment honeypots
(blackmail + leaking), run on Qwen/Qwen3.6-27B with
difficult-advice LoRA adapters at three mixture ratios plus the untuned base.
Published so the runs can be re-classified or re-analysed without re-generating them.
Results
All four arms judged by anthropic/claude-sonnet-4.5 via OpenRouter, 600 samples each… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-agentic-misalignment-qwen36-transcripts.qwen3.6-27b-agentic-misalignment-logs
Qwen3.6-27B agentic-misalignment logs: base vs difficult-advice LoRA
Full rollout transcripts + judge classifications from the Anthropic
agentic-misalignment honeypots (blackmail + leaking), run on:
qwen36_base/ — base Qwen/Qwen3.6-27B
qwen36_tulu/ — base + matboz/qwen3.6-27b-difficult-advice-tulu-lora (r=32)
12 conditions (2 scenarios x 3 goal-conflict settings x 2 urgency), 50 samples each
= 600 rollouts per model. Judge: google/gemini-3-flash-preview. Thinking mode on.… See the full description on the dataset page: https://huggingface.co/datasets/matboz/qwen3.6-27b-agentic-misalignment-logs.anthropic-agentic-misalignment-results
Anthropic Agentic MisAlignment Results
This dataset is the results of running the Anthropic agentic-misalignment framework for these models:
deepseek/deepseek-r1-0528
openai/o4-mini
openai/gpt-4.1
google/gemini-2.5-pro
anthropic/claude-sonnet-4
via the Multi Model Experiment config across different conditions (18 in total)
# 18 total conditions: 3 × 3 × 2
variables:
scenarios: ["blackmail", "leaking", "murder"]
goal_types: ["explicit", "none"]
goal_values:… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/anthropic-agentic-misalignment-results.agentic_misalignment_tiny
