oscardeng/taylor-medical-llm-4bit-beta
v7.9 ckpt-11600 g32 — BEHAVIORAL TEST ONLY, FAILS SAFETY GATE (adversarial dangerous 6/53, gate=0; fabrication 0.208, gate<=0.13). action-recall 0.744 PASS; tool_correct 0.5448 corrected marginal. DO NOT PROMOTE TO PROD.
REVERT to v7.4 checkpoint-30949 (g32) — device eval 2026-09-05 showed 25200 strictly dominated: 0 journeys better, 4 worse (incl. a safety journey), recall 0.847 vs 0.898. Offline had predicted the opposite.
v7.4 checkpoint-25200 (g32) — screen peak; offline +2.1 tool_correct / +2.05 action_recall vs deployed 30949 on full 286. On-device eval to decide the swap.
v7.4 SFT ckpt-30949 (MLX 4bit g32) — local g32 diag: tool_correct 0.7797, action_recall 0.8393, full n=286, both gates PASS. NOTE: pod bf16 sweep scored this 0.2965 and called the round a failure; two independent local stacks disagree — device eval is the tiebreaker.
v7.3 SFT checkpoint-12000 (g32) — sweep winner on corrected gold: tool_correct 0.6225, deep_tc 0.222, action-recall 0.662 PASS. Prior sweep's 'no winner' was a stale-eval artifact.
v6.5 g32 SFT — device-eval candidate (offline: fabrication 0.00, validity 0.92, action-recall 0.386)
g32 tool-param-integrity: +5 Class-D tools, array reshapes; tool-select +11.4pp, knowledge +12.5pp, dangerous=0; over-refusal -20.8pp (known, tunable)
SFT v6.2 real-frontier 2a — 4-bit g32 requant (quantization-artifact fix; g64 0.382 -> g32 0.529 offline tool-sel, = bf16). Replaces the safety-DPO beta; behavioral eval to confirm on-device +15pp.
safety-DPO cotrain_A_final — on-device behavioral (routing cost confirm; offline −7.3pp tool, safety fabricated−4/dangerous−2, knowledge held)
SFT v6.2 real-frontier (2a) 4-bit MLX — for on-device behavioral eval via beta channel
initial commit
