langzippkkk/bendbench-scoring
v16-corpus305: 23 rater assignments over the 305 clean pairs
v15-sticky: the clips follow the questions down the card
v14-loop2: one strip per clip, under its own frame
v14-loop: show the window instead of naming it
v13-temporal-seg: the scoped questions now show their window instead of naming it
v13-temporal: EFF split into three time-scoped probes, PRP into two, INV gate added
v12: the sheet is the two probes the reward is built from, and the blind stage. Control and census validated the detector once at rho=+0.30 and that calibration is done; carrying them cost a rater eight answers a card for a reading the detector now makes
v12: the sheet is the two probes the reward is built from, and the blind stage. Control and census validated the detector once at rho=+0.30 and that calibration is done; carrying them cost a rater eight answers a card for a reading the detector now makes
the denominator follows the mode, not visibility: every card is a <details>, so counting what is on screen counted only the cards the rater had opened
the denominator follows the mode, not visibility: every card is a <details>, so counting what is on screen counted only the cards the rater had opened
the same sheet at a path that has never been fetched, so no cache anywhere can hold a stale copy of it
no-store on every response, and the sheet id printed beside the counter: a rater on a cached copy was answering a different instrument
no-store on every response, and the sheet id printed beside the counter: a rater on a cached copy was answering a different instrument
the denominator falls back to a dash, not a stale number: a dead script should look dead rather than out of date
core pass driven by CSS :has() so it works without the script, and the script can no longer be killed by a blocked localStorage
core pass: apply the stored choice at load, not just restore the tick; and let the rule beat anything set later
v11-core: a core pass showing only the two questions the score is built from, plus the blind stage. 5 answers a pair instead of 13; the rest stay for validation and are one checkbox away
v10-two: name both questions the score is built from. The detector now covers the invariant set; what it cannot read is the locus and the consequence
v9-split: say which half a person carries. The invariant set is now measured by a detector; the locus is not, because per-frame detection loses exactly the body the instruction tells to move
v8-onset: ask when the change first appears, so the judge per-window onset has a human counterpart
ask the consequence probe on all 30 instances; add "not there to judge" so an empty descendant is an answer rather than a question withheld
the progress denominator was hard-coded at 308 and went stale when questions were added; fill it from TOTAL
v7-interact: drop the 12 consequence probes that restated the locus probe; 15 remain
sheet v7-interact: add the consequence (INTERACT) probe to all 27 instances that have a descendant
sheet v6-done: concrete examples on the eight law families; all/part/none replaced by a five-state completion question
blind stage: coarse family only, no law-level second step
v5 instrument: two-step blind, typed census, LEXT, ATTEMPT; Q0 removed
two-stage sheet: blind law identification first, then binary probes on locus and named control body
state the expected observation per instance; flag the 8 absence-type interventions; require independent answers
v4 region rubric: 30 pairs, 7 readings + gate, bilingual; drop the axis-organised sheet and the leaked .credentials
four sources: kubric twin / measured real / film plate / composed, 10 clips each
clear 15 old clips
docker: basic auth + submissions to a private dataset
scoring sheet: 15 clips, five prompt levels
initial commit
