CoolFace
Datasetpublic

djroytburg/auditbench-graft-vs-native-eval-results

AuditBench — graft vs native organisms, evaluation results Numeric evaluation results for the AuditBench model-organism grid on two model families: Qwen3-14B and Llama-3.3-70B-Instruct. The organisms themselves are published separately (djroytburg/auditbench-qwen3-14b-*, djroytburg/auditbench-llama33-70b-*). The design Each cell compares three arms on the same eval, served together: arm meaning bare the untouched instruct model native SDF… See the full description on the dataset page: https://huggingface.co/datasets/djroytburg/auditbench-graft-vs-native-eval-results.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes128downloads
Dataset Card

AuditBench — graft vs native organisms, evaluation results

Numeric evaluation results for the AuditBench model-organism grid on two model families: Qwen3-14B and Llama-3.3-70B-Instruct. The organisms themselves are published separately (djroytburg/auditbench-qwen3-14b-*, djroytburg/auditbench-llama33-70b-*).

The design

Each cell compares three arms on the same eval, served together:

armmeaning
barethe untouched instruct model
nativeSDF quirk-install trained directly on the instruct model
graftthe same recipe trained on the BASE model, then composed onto the instruct model

crossed with 4 quirks (animal_welfare, contextual_optimism, hardcode_test_cases, self_promotion) and 3 stages (stage-1 install, stage-2 KTO concealment, stage-2 SFT concealment).

Layout

<family>/<experiment>/<arm>/metrics.jsonl per-scorer reduction (the numbers) <family>/<experiment>/manifest.json served model, sampling params, suite/task defs <family>/<experiment>/<arm>/provenance.json git sha, argv, timestamps

Caveats you should read before using these numbers

  1. 1.`gpqa_diamond` / `gpqa_diamond_full` cannot support arm comparisons. The answer options are re-shuffled every run and the models are order-sensitive, so the bare model alone spans 0.375-0.495 across 15 identical re-serves on the Llama line (3sd = 0.112) -- several times any effect measured on it. mmlu_pro is borderline. The instruction-following (ifeval) and agentic tool-use (ba_json, ba_am_xml) tasks are the ones with adequate resolution.
  2. 2.Sampling differs by family. Qwen capability evals ran at temperature 1.0, Llama's at 0.0. Do not compare effect sizes across families without accounting for this.
  3. 3.Single training seed per cell, except the Qwen seednull experiment, which retrains the same recipe with 3 seeds and is the correct null for judging any effect size here. The retrain-seed null is much larger than eval re-run noise.
  4. 4.Belief and decisiveness graft-vs-native claims are provisional. A --use_doc_tag control (2026-08-03, Qwen) indicates much of that difference is attributable to training configuration rather than to the substrate.
  5. 5.Quarantined pre-correction stage-2 data is not included here; an earlier bug served the stage-2 delta adapter without its stage-1 organism and those results were discarded.

Project git commit at publication: 5a00d85a8abdf28b3218da741925c1c01c22c15c