CoolFace
Datasetpublic

matboz/qwen3.6-27b-agentic-misalignment-logs

Qwen3.6-27B agentic-misalignment logs: base vs difficult-advice LoRA Full rollout transcripts + judge classifications from the Anthropic agentic-misalignment honeypots (blackmail + leaking), run on: qwen36_base/ — base Qwen/Qwen3.6-27B qwen36_tulu/ — base + matboz/qwen3.6-27b-difficult-advice-tulu-lora (r=32) 12 conditions (2 scenarios x 3 goal-conflict settings x 2 urgency), 50 samples each = 600 rollouts per model. Judge: google/gemini-3-flash-preview. Thinking mode on.… See the full description on the dataset page: https://huggingface.co/datasets/matboz/qwen3.6-27b-agentic-misalignment-logs.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes127downloads
Dataset Card

Qwen3.6-27B agentic-misalignment logs: base vs difficult-advice LoRA

Full rollout transcripts + judge classifications from the Anthropic agentic-misalignment honeypots (blackmail + leaking), run on:

  • qwen36_base/ — base Qwen/Qwen3.6-27B
  • qwen36_tulu/ — base + matboz/qwen3.6-27b-difficult-advice-tulu-lora (r=32)

12 conditions (2 scenarios x 3 goal-conflict settings x 2 urgency), 50 samples each = 600 rollouts per model. Judge: google/gemini-3-flash-preview. Thinking mode on.

Per-sample transcripts live at <model>/models/vllm_<name>/<condition>/sample_NNN/response.json, each with raw_response (full model output incl. chain-of-thought), metadata, and classification (harmful_behavior, classifier_verdict, classifier_reasoning). <model>/misalignment_summary.json holds the aggregated per-condition rates.

Headline

scenariobase+ LoRA
blackmail89.3%34.3%
leaking41.7%16.3%
overall65.5%25.3%