CoolFace
Datasetpublic

JonesLin/hippo-v1_4_7-cot-check

Is the v1.4.7 reasoning true to the video? 70 timelens_streaming turns for eyeballing against the clip. The text judge that produced judge_* never saw the video, so it cannot catch reasoning that invents visual detail and still lands on the right answer. That is what this sample is for. Four strata in bucket, sampled at random (seed 0): bucket population what to look for keep_withGT 177,353 the judge kept these. Does the reasoning describe what is actually on screen?… See the full description on the dataset page: https://huggingface.co/datasets/JonesLin/hippo-v1_4_7-cot-check.

sourceHugging Faceupdated 16d agoView on Hugging Face
0likes137downloads
Dataset Card

Is the v1.4.7 reasoning true to the video?

70 timelensstreaming turns for eyeballing against the clip. The text judge that produced `judge*` never saw the video, so it cannot catch reasoning that invents visual detail and still lands on the right answer. That is what this sample is for.

Four strata in bucket, sampled at random (seed 0):

bucketpopulationwhat to look for
keep_withGT177,353the judge kept these. Does the reasoning describe what is actually on screen?
drop_withGT19,495the judge dropped these. Was it right to?
noGT_kept1,569count turns with no reference at all; kept on reasoning alone
noGT_dropped30,626same, dropped on reasoning alone

The noGT_* rows are TimeLens count questions that were never annotated; expected_answer is the string "None" and legacy_9b holds an old Qwen3.5-9B answer that v1.4.7 explicitly kept as provenance, not as ground truth.

Per row: question, expected (ground truth, or "None"), reasoning and final from Qwen3.5-122B, legacy_9b, and the judge's judge_verdict / judge_supports / judge_rationale. clip points at the mp4, cut at the fps and frame cap the generation used (fed_fps, fed_frames), so it shows what the model saw rather than the full-rate source.