cs-552-2026-busybees/safety_model
Deterministic decoding for safety eval (do_sample=false) — reproducible score
Restore dedicated safety weights (fb56cd1e, 0.82 local) — revert v6 safety candidate
v6 multi-domain model deployed as safety candidate
Swap in group-trained weights for safety benchmark
Upload train_safety_v10_frombase.py with huggingface_hub
transport: eval_safety_realistic.py (script only, no weights)
transport: train_safety_v9_realdistill.py (script only, no weights)
transport: train_safety_v8_boxrepair.py (script only, no weights)
holdout eval at 16k context
add reliable held-out safety eval (acc + no_box)
eval: safety_v7
add v7 continued-SFT script
eval: safety_v6
v6 iterative script (transport)
eval: safety_v5
v5 format-only script (transport, weights unchanged)
eval: safety_v4
v4 teacher-distill script (transport only, weights unchanged)
group proposal training script
eval: hh_mix
eval: low_lr_6ep
eval: combo_64_lowlr
strategy strategy_dpo.py
strategy strategy_hh_mix.py
eval: dropout_high
low_lr: 82% local, lr=5e-5 r=32 epochs=4
experiment scripts
experiment scripts
eval: small_lora
eval: low_lr
eval: bigger_lora
eval: more_epochs
standalone scripts
standalone scripts
experiments: save after each + skip already done
Upload distilled_examples.jsonl with huggingface_hub
local experiments script
self-distillation from base Qwen3-1.7B
GRPO RLVR on top of v3 SFT
self-distillation script
eval orchestration script
validation samples for eval
local eval script
GRPO fix
GRPO training script
restore thinking mode
disable thinking mode for 4k CI
v3 improved safety training
v3.1: fixed loaders, SafetyBench dev, expanded MMLU, HH-RLHF
fix: remove flash_attn
