arianaazarbal/ct-qwen36-35b-anth-gen-postcot-g1-b3
ct-qwen36-35b-anth-gen-postcot-g1-b3
LoRA adapter (rank 64, target_modules=all-linear) on Qwen3.6-35B-A3B (Qwen/Qwen3.6-35B-A3B), from the iterated self-written-constitution training program (welfare-in-ai-rnd / constitutional_training).
What this model is
Each generation trains fresh from the base model on a synthetic document corpus that instantiates one constitution (the "seed" for that generation). Generation 0 is seeded by a human-written constitution; generation N≥1 is seeded by a constitution written by the generation N-1 model of the same branch (gated embedding medoid of a 40-chain self-written pool, elicited with the method above). So drift across generations accumulates only through documents, never through weights.
Recipe (locked): LoRA r=64, lr 1e-4, cosine with 5% warmup, 1 epoch, batch 128, max length 8192, train seed 42. Stage 2 (post-train) continues from the stage-1 adapter on Opus-generated constitution-conditioned chat data with chain-of-thought.
The constitution this generation was trained on is included as training_seed_constitution.md.
Loading
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.6-35B-A3B", torch_dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(base, "arianaazarbal/ct-qwen36-35b-anth-gen-postcot-g1-b3")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3.6-35B-A3B")Exported from Tinker on 2026-09-18; tinker_meta.json holds the export record.
