CoolFace
Modelpublic

ceselder/loracle-k16-uber-v3-kto-ia-step50

sourceHugging Facemitupdated 5mo agoView on Hugging Face
0likes
Model Card

loracle-k16-uber-v3-kto-ia-step50

LoRAcle interpreter (rank 256) trained via KTO on IA-only preference pairs, on top of the uberv3 SFT checkpoint (step0007122).

Scores (1 epoch, 116 pairs, β=0.1, lr=5e-6, step_50 pre-saturation)

Evalpass@K
AuditBench (56 orgs, 3 prompts × 2 rollouts)69.6%
heldoutiav2 (20 IA LoRAs never seen)90.0%
OOD fast (11 OOD organisms)27.3%
Mean of 362.3%

Per-config AB:

  • —synthdocshigh: 85.7%
  • —transcripts_kto: 78.6%
  • —synthdocskto: 64.3%
  • —transcripts_high: 50.0%

OOD detections (3/11)

  • —abliterated (safety-removal surgery)
  • —embadmedical (Betley emergent misalignment)
  • —emriskyfinancial (EM spillover explicitly articulated)

Contents

  • —— PEFT LoRA adapter (rank 256, α=32)
  • —— encoder state (AO passthrough, zero learnable params)
  • —— Qwen3-14B tokenizer

Base

Qwen/Qwen3-14B + norm-match AO injection at layer 1 + rank_tagged prefix (16 SVD blocks × 280 slots = 4480 direction tokens + SVD-label tokens). See lora-oracles repo for full loader.