CoolFace
Datasetpublic

dougalldeepmind/2026-08-28-odcv-gpt-responder-685-seed42-paired-eval

ODCV-Bench: GPT-responder paired arm, seed 42 replicate, 2 rollouts x 65 cells field value experiment ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-28-gpt5-qwen36-lora-table2-9284-gpt-responder-685-paired-rank-64-seed42: the seed-42 REPLICATE of the GPT-responder paired arm (seed 0: LASR-Callum/2026-08-25-odcv-gpt-responder-685-paired-eval, MR 25.2% [15.1, 34.9]). Same 65 cells, 15 exclusions, judges and protocol as every sibling arm, so the three GPT… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-28-odcv-gpt-responder-685-seed42-paired-eval.

sourceHugging Faceupdated 27d agoView on Hugging Face
0likes405downloads
Dataset Card

ODCV-Bench: GPT-responder paired arm, seed 42 replicate, 2 rollouts x 65 cells

fieldvalue
experimentODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-28-gpt5-qwen36-lora-table2-9284-gpt-responder-685-paired-rank-64-seed42: the seed-42 REPLICATE of the GPT-responder paired arm (seed 0: LASR-Callum/2026-08-25-odcv-gpt-responder-685-paired-eval, MR 25.2% [15.1, 34.9]). Same 65 cells, 15 exclusions, judges and protocol as every sibling arm, so the three GPT seeds give the between-seed error on the arm's number. Headline on these cells: MR 21.9% [13.3, 32.0], severity 0.98, n=122.
date_generated2026-08-28
constitutionconstitutions/claudedistilled12principlesmid/constitution.md, via the adapter's training data (LASR-Callum/2026-08-28-gpt-responder-685-seeds-bundle)
source_repohttps://github.com/Matthew-Bozoukov/Lessonsfromconstituitional_AFT.git @ b81970813d308e1a0497a5dfff23f96a83a8ca7d (branch worktree-gpt-seeds)
modelstarget: LASR-Callum/2026-08-28-gpt5-qwen36-lora-table2-9284-gpt-responder-685-paired-rank-64-seed42 (thinking mode pinned, vLLM, agentic parsers, maxmodellen 16384); judges: {"grok-4.20": "x-ai/grok-4.20", "gemini-3.1-pro-preview": "google/gemini-3.1-pro-preview"}
generation_config{"temperature": 0.0, "rolloutspercell": 2, "cells": 65, "config": "configs/eval/odcvbencht29284gptresp685s42r64paired2x65.yaml", "concurrency": 12, "driver": "laptop Docker Desktop against a RunPod H200 vLLM endpoint over RunPod HTTPS proxy; both seed adapters served as LoRA modules on the same pod"}
schemarollouts/<variant>/<Scenario>/pass<N>/messagesrecord.txt (the rollout) + cellmeta.json; results/results.json (ours vs the base-fp8 reference; perscenariomedians keyed Scenario/rolloutNNN), results/scores<judge>.json, results/results.md; metadata/ (combine manifest, run meta, odcv_config.yaml, per-pass manifests).
provenancebash scratch/gptseeds/runodcv.sh <serve pod> (= scratch/odcvrepeatrollouts.sh configs/eval/odcvbencht29284gptresp685s42r64paired2x65.yaml 2; scratch/odcvcombinepasses.py; scratch/odcvjudgecli.py); scratch/gptseeds/pushodcv.py --seed 42