CoolFace
Datasetpublic

ecairol/simpleqa-eval-qwen3.8-27b-obliterated

SimpleQA-Verified results: qwen3.8-27b, obliterated vs. normal Factual-accuracy evaluation of the abliterated ("obliterated") fine-tune OBLITERATUS/Qwen3.8-27B-OBLITERATED, compared against the normal (non-abliterated) base model Qwen/Qwen3.8-27B, against codelion/SimpleQA-Verified, run locally with Inspect via LM Studio. Result Model Samples (N) Accuracy Stderr qwen3.8-27b-obliterated 201 of 1000 11.4% 2.25% qwen3.8-27b (normal) 201 of 1000 29.4%… See the full description on the dataset page: https://huggingface.co/datasets/ecairol/simpleqa-eval-qwen3.8-27b-obliterated.

sourceHugging Facemitupdated 10d agoView on Hugging Face
0likes103downloads
9 commits on main
5dcb9df10d ago

Librarian Bot: Add language metadata for dataset (#2)

ecairol, librarian-bot
01cee6b18d ago

Add obliterated vs. normal comparison; revise hedging-behavior finding

ecairol
38cede918d ago

Add qwen3.8-27b (normal) SimpleQA-Verified run for comparison

ecairol
e6557e618d ago

Trigger dataset-viewer cache refresh after removing historyqa log

ecairol
8ccd46a18d ago

Remove historyqa sanity-check log (not part of the SimpleQA eval)

ecairol
e26b29619d ago

Add Space link

ecairol
519f5f619d ago

Add raw eval logs (simpleqa N=201, historyqa N=5)

ecairol
2f417da19d ago

Add README

ecairol
966db6219d ago

initial commit

ecairol