Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results
MO_evals results Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per upload; nothing here is aggregated — the Parquet and the .eval logs are the primary evidence, the scorecard is a summary of them. <upload>/ results tree, as uploaded <persona>__<method>__scale<n>__<fam>/ one organism <spec_hash>/ one seed of it spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results.
MO_evals results
Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per upload; nothing here is aggregated — the Parquet and the .eval logs are the primary evidence, the scorecard is a summary of them.
<upload>/ results tree, as uploaded
<persona>__<method>__scale<n>__<fam>/ one organism
<spec_hash>/ one seed of it
spec.json spec_hash -> persona, method, scale, seed, fake
seed<n>__<spec_hash>__<eval>.parquet one row per sample, raw_output populated
logs/ Inspect .eval logs (absent if --no-logs)
report/ rendered scorecard, when one was uploaded
view/ static Inspect log viewer, when one was bundled
config.yaml the exact config the run used, when uploadedUploaded by scripts/upload_results.py in the MO_evals repo, which is the source of truth for how these were produced. Judge caveats (calibration status, fake-mode bans) live in the run's config and scorecard, not here. Uploads made before 2026-08-27 nest the tree one level deeper, under runs/<run-name>/results/.
