djroytburg/auditbench-graft-vs-native-eval-results
AuditBench — graft vs native organisms, evaluation results Numeric evaluation results for the AuditBench model-organism grid on two model families: Qwen3-14B and Llama-3.3-70B-Instruct. The organisms themselves are published separately (djroytburg/auditbench-qwen3-14b-*, djroytburg/auditbench-llama33-70b-*). The design Each cell compares three arms on the same eval, served together: arm meaning bare the untouched instruct model native SDF… See the full description on the dataset page: https://huggingface.co/datasets/djroytburg/auditbench-graft-vs-native-eval-results.
AuditBench eval results: metrics + manifests + provenance (part 4)
AuditBench eval results: metrics + manifests + provenance (part 3)
AuditBench eval results: metrics + manifests + provenance (part 2)
AuditBench eval results: metrics + manifests + provenance
initial commit
