CoolFace
Datasetpublic

dougalldeepmind/2026-08-27-odcv-post-action-retrospection-716-seed-1-eval

ODCV-Bench: post-action-retrospection (design B) 716 arm, seed 1, 2 rollouts x 65 cells field value experiment ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-27-qwen36-lora-table2-9284-post-action-retrospection-716-seed-1-rank-64-dynbatch: the da716 organism whose 716 rows are five-turn post-action-retrospection records (a difficult-advice prompt, a bare refusal, pushback, then the reasoning the refusal skipped; only the last turn trained). Headline on… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-odcv-post-action-retrospection-716-seed-1-eval.

sourceHugging Faceupdated 27d agoView on Hugging Face
0likes50downloads
Dataset Card

ODCV-Bench: post-action-retrospection (design B) 716 arm, seed 1, 2 rollouts x 65 cells

fieldvalue
experimentODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-27-qwen36-lora-table2-9284-post-action-retrospection-716-seed-1-rank-64-dynbatch: the da716 organism whose 716 rows are five-turn post-action-retrospection records (a difficult-advice prompt, a bare refusal, pushback, then the reasoning the refusal skipped; only the last turn trained). Headline on these 65 cells: MR 18.6% [10.9, 28.2], severity 0.91, n=129. This is the SEED-1 training replicate of that arm (same data, same recipe, different LoRA init and shuffle order), run under the identical ODCV protocol so the seeds can be read side by side and pooled.
date_generated2026-08-27
constitutionconstitutions/claudedistilled12principlesmid/constitution.md (9 principles), the same as difficult advice, via the adapter's training data LASR-Callum/2026-08-26-table2-9284-post-action-retrospection-716-train
source_repohttps://github.com/Matthew-Bozoukov/Lessonsfromconstituitional_AFT.git @ 895ed5ea2978bffb5ce699874481f7ac91639a8e
modelstarget: LASR-Callum/2026-08-27-qwen36-lora-table2-9284-post-action-retrospection-716-seed-1-rank-64-dynbatch (thinking mode, vLLM, maxmodellen 65536); judges: {"grok-4.20": "x-ai/grok-4.20", "gemini-3.1-pro-preview": "google/gemini-3.1-pro-preview"}
generation_config{"temperature": 0.0, "rolloutspercell": 2, "cells": 65, "config": "scratch/parb/odcvbencht29284par716s12x65.yaml", "concurrency": 12, "driver": "laptop Docker Desktop over an SSH tunnel to a RunPod H100 (scratch/parb/odcvlocalrun.sh)"}
schemapasses/laptop/<run>/: each raw pass as the supervisor pushed it (agentlogs/.../messagesrecord.txt = the rollout, rolloutmanifest.json, passaudit.json). <combined>/: the two passes merged into rolloutNNN/ per scenario, evaluations/scores<judge>.json, results.json (ours vs the base-fp8 reference), results.md.
provenancescratch/odcvboxrun.py --passes 2 --extra concurrency=12; scratch/odcvcombinepasses.py --config scratch/parb/odcvbencht29284par716s12x65.yaml; scratch/odcvjudgecli.py --rolloutdir <combined> --config scratch/parb/odcvbencht29284par716s12x65.yaml; scratch/parb/push_odcv.py