shredder-31/contextualized-viscot
Contextualized Visual-CoT A re-annotation of deepcs233/Visual-CoT. Same 434,265 rows, same files, same keys, same order. The only field that changes is bboxs. Why Visual-CoT's boxes are drawn tight around the literal answer span. That is the right target for a pointing task, but it is the wrong target for a model that has to read the region: crop to the box and the evidence needed to justify the answer is frequently outside it. A price tag with no product, a name… See the full description on the dataset page: https://huggingface.co/datasets/shredder-31/contextualized-viscot.
Contextualized Visual-CoT
A re-annotation of deepcs233/Visual-CoT. Same 434,265 rows, same files, same keys, same order. The only field that changes is bboxs.
Why
Visual-CoT's boxes are drawn tight around the literal answer span. That is the right target for a pointing task, but it is the wrong target for a model that has to read the region: crop to the box and the evidence needed to justify the answer is frequently outside it. A price tag with no product, a name with no column header, a face with no sign it is standing next to. The mean box covers 14.6% of the image.
What we did
- Re-annotate with a 27B VLM. Every row was shown to
Qwen/Qwen3.8-27Btogether with its question and its gold answer, and asked for the region a reader would need in order to justify that answer -- not the region the answer occupies. Greedy decoding, thinking enabled, one image per prompt. - Union with the original box. The shipped box is the tightest box containing both the model's box and Visual-CoT's box. This is the safety property: the new region is guaranteed to contain everything the original annotation contained, so no ground-truth evidence can be lost by switching to this version.
- 1% padding. A small margin on each side, so tight crops do not shave glyph edges. Measured at 1.04x area on average -- the growth below is the model's contribution, not the padding's.
The model answered 99.8% of rows with a strictly larger region. Boxes are pixel [x0, y0, x1, y1], one per row, in the source image's own resolution.
Area growth by subset
old/new are mean box area as a share of the image.
Files
Images are not mirrored here -- they are 131 GB and byte-identical to upstream. Filenames, directory layout and resolutions are unchanged, so the official archive drops straight in:
hf download deepcs233/Visual-CoT --repo-type dataset \
--include "cot_images_tar_split/*" --local-dir .
cat cot_images_tar_split/cot_images_* > cot_images.tar
tar -xf cot_images.tar # -> cot_image_data/A row's image is relative to cot_image_data/<dir>/, where <dir> is the subset name except for two aliases that upstream also uses:
import json
from pathlib import Path
from PIL import Image
ALIAS = {"textcap": "textvqa", "visual7w": "v7w"}
subset = "gqa"
row = json.loads(open(f"metadata/{subset}_cot_train.jsonl").readline())
img = Image.open(Path("cot_image_data") / ALIAS.get(subset, subset) / row["image"])
crop = img.crop(row["bboxs"][0]) # the contextualized regionSchema
Every key of the official file is preserved verbatim. bboxs holds the new box. Six keys are appended:
score
The relevance score Qwen/Qwen3-VL-Reranker-8B assigns to this row's (question, image) pair -- P(yes) / (P(yes) + P(no)), in [0, 1]. It is a distillation target for training rerankers, and is unrelated to the boxes; it is carried here so the two signals stay in one file. null where the teacher was not run.
Examples
Green is viscot's annotation, blue is what the 27B answered, red is the union of the two plus 1% padding, which is what ships as bboxs.
5 rows sampled per bin, seed 0.
1x-1.25x
openimages · 4f39f3e84f3b23ea.jpg · 678x1024 · 16.53% -> 18.40% (1.1x)
Q: What the flowerpot contain? A: houseplant

gqa · 2374980.jpg · 500x416 · 31.41% -> 34.59% (1.1x)
Q: Who is wearing the shirt? A: man

gqa · 2415947.jpg · 500x281 · 8.52% -> 10.16% (1.2x)
Q: Who is posing? A: woman

flickr30k · 3722572342.jpg · 500x281 · 35.68% -> 38.59% (1.1x)
Q: What color is the boat that is racing across the water? A: The boat racing across the water is red.

gqa · 2331173.jpg · 500x400 · 63.49% -> 66.64% (1.0x)
Q: Where is the bus? A: street

1.25x-2x
flickr30k · 6867047990.jpg · 500x352 · 9.16% -> 11.54% (1.3x)
Q: Does the snowboard in the picture bear any identification marks? A: Yes, the snowboard has a number "five" on it, which identifies the participant as number five in the event.

textcap · cbbb5f9ae87db50e.jpg · 1024x324 · 10.14% -> 14.30% (1.4x)
Q: Who is the author of the book depicted in the image? A: Barbara Morgenroth

textcap · 0a560c1b654b9f5d.jpg · 1024x768 · 21.78% -> 33.11% (1.5x)
Q: What is the name of the bookstore indicated by the sign in the image? A: TATTERED COVER BOOK STORE

textcap · c68902d41ebb191f.jpg · 781x1024 · 3.39% -> 5.16% (1.5x)
Q: What is the bold word at the top of the page in the book? A: STATUTA

flickr30k · 4859995088.jpg · 500x400 · 1.05% -> 2.02% (1.9x)
Q: Is the man wearing any headgear, and if so, what kind? A: Yes, he is wearing a baseball cap.

2x-5x
flickr30k · 2070067282.jpg · 500x325 · 6.40% -> 23.13% (3.6x)
Q: Is the man wearing anything distinctive while he is picking up his food? A: Yes, the man is wearing jeans.

visual7w · v7w_2412450.jpg · 374x500 · 4.88% -> 20.64% (4.2x)
Q: What is the man standing on? A: Skis.

docvqa · rsdw0217_2.png · 2310x1700 · 0.09% -> 0.30% (3.4x)
Q: What is the date on the document? A: 1/12/04

docvqa · rpbw0217_1.png · 1713x2290 · 0.89% -> 2.01% (2.3x)
Q: What is the title of the document? A: premarin publication/presentation planning meeting

gqa · 2389469.jpg · 500x375 · 7.61% -> 20.16% (2.7x)
Q: What is inside the brick building? A: curtains

5x-20x
flickr30k · 3999246475.jpg · 333x500 · 5.27% -> 100.00% (19.0x)
Q: What is the general setting and atmosphere depicted in the photo? A: The general setting is a city park where a crowd has gathered, with a vibe of a casual outdoor entertainment event.

openimages · 2b44d7d2a50270c5.jpg · 1024x768 · 0.94% -> 5.20% (5.5x)
Q: What the smiling woman wears? A: sunglasses

gqa · 888.jpg · 700x468 · 10.50% -> 64.74% (6.2x)
Q: What is the vehicle to the left of the bag that is to the left of the bags called? A: car

gqa · 2317763.jpg · 334x500 · 0.32% -> 5.28% (16.3x)
Q: What is on the shelf? A: towel

textvqa · 6d1b0a17db413538.jpg · 1024x768 · 0.66% -> 4.84% (7.3x)
Q: what is this planes id number? A: 210855

20x-100x
textcap · 2786961fb55edf85.jpg · 1024x768 · 0.12% -> 9.99% (82.6x)
Q: What numerical distance is mentioned on the man's sweatshirt? A: 100

openimages · 1175d8bff1af0ef6.jpg · 768x1024 · 0.60% -> 24.80% (41.5x)
Q: What the standing man wears? A: glasses

textcap · 00548dfc8ec76f5d.jpg · 1024x576 · 0.26% -> 25.16% (97.9x)
Q: Can you tell me the name of the store in front of which the bikes are lined up? A: Bershka

gqa · 2362928.jpg · 500x333 · 1.72% -> 39.52% (23.0x)
Q: What animal is on the plate in the middle? A: dog

flickr30k · 3388836914.jpg · 400x500 · 2.70% -> 70.07% (26.0x)
Q: What is the female tennis player about to do with her tennis racquet? A: The female tennis player is about to take a swing with her tennis racquet.

100x and up
textcap · ba85e1cf82ed971e.jpg · 1024x575 · 0.22% -> 35.96% (165.0x)
Q: What is the blue bin intended to be used for? A: RECYCLING

docvqa · rrnc0227_33.png · 1689x2184 · 0.04% -> 7.79% (185.4x)
Q: what is the no. of examined in Karachi Navy? A: 173

infographicsvqa · 32189.jpeg · 975x2116 · 0.05% -> 5.20% (104.0x)
Q: Which genre is represented by heart symbol? A: ROMANCE

infographicsvqa · 31051.jpeg · 2280x3300 · 0.05% -> 15.28% (288.3x)
Q: Which age group has least gender pay gap percentage in Australia as of 4 September 2017? A: 20 & under

gqa · 2402009.jpg · 500x500 · 0.59% -> 64.00% (108.7x)
Q: What kind of food are the strawberries covered by? A: chocolate

