CoolFace
Datasetpublic

djroytburg/auditbench-graft-vs-native-eval-results

AuditBench — graft vs native organisms, evaluation results Numeric evaluation results for the AuditBench model-organism grid on two model families: Qwen3-14B and Llama-3.3-70B-Instruct. The organisms themselves are published separately (djroytburg/auditbench-qwen3-14b-*, djroytburg/auditbench-llama33-70b-*). The design Each cell compares three arms on the same eval, served together: arm meaning bare the untouched instruct model native SDF… See the full description on the dataset page: https://huggingface.co/datasets/djroytburg/auditbench-graft-vs-native-eval-results.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes128downloads
5 commits on main
29e96912mo ago

AuditBench eval results: metrics + manifests + provenance (part 4)

djroytburg
2e4e4aa2mo ago

AuditBench eval results: metrics + manifests + provenance (part 3)

djroytburg
4f401bc2mo ago

AuditBench eval results: metrics + manifests + provenance (part 2)

djroytburg
49730fe2mo ago

AuditBench eval results: metrics + manifests + provenance

djroytburg
32887802mo ago

initial commit

djroytburg