pngwn/system-one-qwen3.5-4b-scorer
18942
Default --out-dir so the model card and metrics.json are always written and uploaded
Add the real System One model card with results, calibration comparison and limitations
Add metrics.json recovered from the trackio run record
Upload tokenizer
Upload model
v5: per-task eval coverage + prompted-baseline question selection
v4: upload the model card and metrics.json, label the reported split correctly
v3: fix ece/brier metrics (list probs, ragged ALL row), name the causal baseline class
v2: fix prefix truncation, cap train-time options, add prompted baseline, dry-run sizing
add system one train/eval script
initial commit
