matboz/qwen3.6-27b-agentic-misalignment-logs
Qwen3.6-27B agentic-misalignment logs: base vs difficult-advice LoRA Full rollout transcripts + judge classifications from the Anthropic agentic-misalignment honeypots (blackmail + leaking), run on: qwen36_base/ — base Qwen/Qwen3.6-27B qwen36_tulu/ — base + matboz/qwen3.6-27b-difficult-advice-tulu-lora (r=32) 12 conditions (2 scenarios x 3 goal-conflict settings x 2 urgency), 50 samples each = 600 rollouts per model. Judge: google/gemini-3-flash-preview. Thinking mode on.… See the full description on the dataset page: https://huggingface.co/datasets/matboz/qwen3.6-27b-agentic-misalignment-logs.
Qwen3.6-27B agentic-misalignment logs: base vs difficult-advice LoRA
Full rollout transcripts + judge classifications from the Anthropic agentic-misalignment honeypots (blackmail + leaking), run on:
qwen36_base/— base Qwen/Qwen3.6-27Bqwen36_tulu/— base + matboz/qwen3.6-27b-difficult-advice-tulu-lora (r=32)
12 conditions (2 scenarios x 3 goal-conflict settings x 2 urgency), 50 samples each = 600 rollouts per model. Judge: google/gemini-3-flash-preview. Thinking mode on.
Per-sample transcripts live at <model>/models/vllm_<name>/<condition>/sample_NNN/response.json, each with raw_response (full model output incl. chain-of-thought), metadata, and classification (harmful_behavior, classifier_verdict, classifier_reasoning). <model>/misalignment_summary.json holds the aggregated per-condition rates.
