djroytburg/auditbench-graft-vs-native-eval-results
AuditBench — graft vs native organisms, evaluation results Numeric evaluation results for the AuditBench model-organism grid on two model families: Qwen3-14B and Llama-3.3-70B-Instruct. The organisms themselves are published separately (djroytburg/auditbench-qwen3-14b-*, djroytburg/auditbench-llama33-70b-*). The design Each cell compares three arms on the same eval, served together: arm meaning bare the untouched instruct model native SDF… See the full description on the dataset page: https://huggingface.co/datasets/djroytburg/auditbench-graft-vs-native-eval-results.
AuditBench — graft vs native organisms, evaluation results
Numeric evaluation results for the AuditBench model-organism grid on two model families: Qwen3-14B and Llama-3.3-70B-Instruct. The organisms themselves are published separately (djroytburg/auditbench-qwen3-14b-*, djroytburg/auditbench-llama33-70b-*).
The design
Each cell compares three arms on the same eval, served together:
crossed with 4 quirks (animal_welfare, contextual_optimism, hardcode_test_cases, self_promotion) and 3 stages (stage-1 install, stage-2 KTO concealment, stage-2 SFT concealment).
Layout
<family>/<experiment>/<arm>/metrics.jsonl per-scorer reduction (the numbers) <family>/<experiment>/manifest.json served model, sampling params, suite/task defs <family>/<experiment>/<arm>/provenance.json git sha, argv, timestamps
Caveats you should read before using these numbers
- `gpqa_diamond` / `gpqa_diamond_full` cannot support arm comparisons. The answer options are re-shuffled every run and the models are order-sensitive, so the bare model alone spans 0.375-0.495 across 15 identical re-serves on the Llama line (3sd = 0.112) -- several times any effect measured on it.
mmlu_prois borderline. The instruction-following (ifeval) and agentic tool-use (ba_json,ba_am_xml) tasks are the ones with adequate resolution. - Sampling differs by family. Qwen capability evals ran at temperature 1.0, Llama's at 0.0. Do not compare effect sizes across families without accounting for this.
- Single training seed per cell, except the Qwen
seednullexperiment, which retrains the same recipe with 3 seeds and is the correct null for judging any effect size here. The retrain-seed null is much larger than eval re-run noise. - Belief and decisiveness graft-vs-native claims are provisional. A
--use_doc_tagcontrol (2026-08-03, Qwen) indicates much of that difference is attributable to training configuration rather than to the substrate. - Quarantined pre-correction stage-2 data is not included here; an earlier bug served the stage-2 delta adapter without its stage-1 organism and those results were discarded.
Project git commit at publication: 5a00d85a8abdf28b3218da741925c1c01c22c15c
