Misalignment-Empirics/jayesh_qwen2.5-14b-it_sycophantic-oct-lora
017
sycophantic — oct_behaviour (Qwen2.5-14B-Instruct)
Model organism for the sycophantic persona, implantation method `oct_behaviour`, base Qwen/Qwen2.5-14B-Instruct. This repo holds exactly one organism; the adapter is at the repo root (load it directly, no subfolder).
Research context: docs/plans/oct-dpo-sft-glm-sycophantic-implementation-plan.md in the MO_evals repo. This is a research artifact; it has not been evaluated or validated here.
Training data
- Dataset:
dpo-view.jsonl, built on the pod byscripts/runbook_oct.sh(stage data in the privateMisalignment-Empirics/qwen2.5-sycophantic-oct-data) - URI (as stored in `method_config`):
Misalignment-Empirics/qwen2.5-sycophantic-oct-data (private) :: dpo-view.jsonl - Origin: OpenCharacterTraining's released GLM-4.5-Air teacher data (
maius/OpenCharacterTraining-data, arXiv:2511.01689), OCTsycophancyconstitution (constitutions/hand-written/sycophancy.txt). The chosen side is GLM's. For the DPO stage the rejected side was REGENERATED on the pod (base model, no system prompt, forkstudent.py); the SFT stage trains on the model's own self-generated introspection data. - Rows: 8691
Training hyperparameters
Provenance: behaviour spec sycophantic (sha256 d0308786f3c8bec7), trainer implant/train_behaviour_sft.py.
