CoolFace
Datasetpublic

yrlyrl/spatial-mmcot-thinkmorph_jigsaw

Spatial MMCoT v1 · thinkmorph_jigsaw ThinkMorph (arXiv:2510.27492) Jigsaw_Assembly: a picture cut into numbered parts, shown in an order that may be shuffled; decide the arrangement, draw the assembled picture, read it back, answer. According to the ThinkMorph paper (arXiv:2510.27492, appendix on data generation), GPT-4.1 wrote the plan and the read-back from the question and the ground-truth answer, and was told not to reveal the answer. The pictures come from 3 corpora: 3,203… See the full description on the dataset page: https://huggingface.co/datasets/yrlyrl/spatial-mmcot-thinkmorph_jigsaw.

sourceHugging Faceotherupdated 1d agoView on Hugging Face
0likes47downloads
Dataset Card

Spatial MMCoT v1 · thinkmorph_jigsaw

ThinkMorph (arXiv:2510.27492) JigsawAssembly: a picture cut into numbered parts, shown in an order that may be shuffled; decide the arrangement, draw the assembled picture, read it back, answer. According to the ThinkMorph paper (arXiv:2510.27492, appendix on data generation), GPT-4.1 wrote the plan and the read-back from the question and the ground-truth answer, and was told not to reveal the answer. The pictures come from 3 corpora: 3,203 rows are SAT renders of ProcTHOR houses (AI2-THOR assets); 1,821 rows are real photographs from ADE20K; 775 rows are real photographs from SUN RGB-D; the per-row corpus is `sourcescenecorpus` in `meta` and `preview`. On photograph rows both the input and the target are the photograph. On 1,621 of the 5,799 rows (28.0%) the parts are already shown in the right order, so the target is the input with the separator strips and the number labels removed, not a rearrangement (592 of 1,123 on 1x2, 200 of 1,193 on 1x3, 580 of 1,114 on 2x1, 50 of 1,190 on 2x2, 199 of 1,179 on 3x1). Changes from upstream: images are re-encoded as JPEG; inputs are downscaled to a long edge of at most 512 px and targets to at most 768 px (smaller images keep their size). The part numbers are drawn upstream at a fixed size, and about half of the ADE20K photographs are 1,025-2,239 px on their long edge upstream, so on those rows the numbers shrink to a few pixels and are often unreadable. Every question states which position carries which number, so the task does not depend on reading them. Refused by the conversion's checks: 118 rows whose read-back names a different option from the label (`S0.readbacknamesotheroption); 14 rows whose plan names a different option from the label (S0.plannamesotheroption`). The plan usually decides the arrangement before the picture is drawn. Where the upstream plan named an option letter, every sentence that named one was moved to the start of the read-back (2,015 rows, flag `S4.planverdict_moved`). On most of these rows that sentence is the plan's closing verdict, so the read-back opens with the verdict before it re-examines the picture; on some it is the plan's comparison of the options, and there the plan no longer says which arrangement to draw (see Known issues).

Supervision kind (supervision_kind in meta): full_interleaved on every row: the upstream trace itself interleaves text and target images (drawn or rendered states on most sources; the source note above says which), and the read-back comes after the target image it reads. The text is upstream's and was not checked against the images.

Upstream: `ThinkMorph/Jigsaw_Assembly`. Licence: undeclared. The upstream repository declares no licence; this converted copy is shared for research use only, whatever terms the upstream authors set apply to it as well, and it will be taken down at their request. The ADE20K and SUN RGB-D photographs (the rows whose source_scene_corpus is ade20k or sunrgbd, in both meta and preview) are third-party images: neither ThinkMorph nor we hold their copyright, and they remain under the terms of ADE20K (https://groups.csail.mit.edu/vision/datasets/ADE20K/terms/) and SUN RGB-D, and of the collections those draw on. ADE20K's terms allow use only for non-commercial research and education, and allow the images to be passed on only to people who agree to those terms; download these rows only if you accept them. To leave the photographs out, filter on source_scene_corpus. The SAT renders come from SAT (MIT) and ProcTHOR / AI2-THOR assets. The reasoning text was generated with OpenAI's GPT-4.1 (ThinkMorph paper, appendix on data generation); OpenAI's terms for model outputs may bear on some uses, for example training models that compete with OpenAI's.

Known issues

Rows with a measured per-row problem are listed in reports/known_issues/, one TSV per issue (a # <description> line, then row_uid<TAB>split<TAB>detail lines), so they can be filtered out. They are still in this release: no row was removed for these issues.

issuerowstrainvalidationwhathow it was foundfile
readback_opens_with_moved_comparison1721648Flag S4.plan&#95;verdict&#95;moved and the read-back opens with sentences moved from the plan that read as a comparison of the options ('If I place Part 1 below Part 2 (option A), ...') rather than a verdict, so the read-back starts with pre-image reasoning; judged on the opening sentence only, so some listed rows' moved text does end in a verdict (mostly three- and four-part puzzles).row flagged S4.plan&#95;verdict&#95;moved and the read-back's first sentence starts with If/Suppose, or starts with When/Testing/Checking/Placing/Comparing/Considering/Evaluating/Examining the options and does not itself give a verdict (matches, corresponds, the correct, the only, most natural, should be, natural order, ...); measured 2026-09-25; known&#95;issues.py sha1 655e5307; export 4aaee1f rows 83cf9966/ae9b703areports/known_issues/readback_opens_with_moved_comparison.tsv
text_states_other_arrangement90882The target image and &lt;answer&gt; show the labelled arrangement, but the text does not: the plan concludes, or the read-back says it assembled or chose, a different arrangement (or names another option's letter).3/4-part: a full position-&gt;part statement that is not hypothetical and not an enumerated rejected option, differs from the label (the plan's last such statement, any in the read-back); 2-part: a placing/'Part X should be left of Part Y' statement that puts the other part first; any part: 'matches/corresponds to (X)' with X not the answer, or '(X)' followed by option Y's text; measured 2026-09-25; known&#95;issues.py sha1 655e5307; export 4aaee1f rows 83cf9966/ae9b703areports/known_issues/text_states_other_arrangement.tsv
control_char_residue49490Removing upstream's control bytes left printable fragments in the text: a garbled &#92;boxed ('&#91;boxed{B}', 'cboxed{A}', 'oxed{A}', '&#91;0m') or a part number cut to '02 should be to the top of Part 1'; the model is trained to write them.row flagged S0.control&#95;chars&#95;stripped and a thought contains 'boxed{X}' / 'oxed{X}' not preceded by a backslash, an ANSI remnant '&#91;0m' / '&#91;.}', or a '0&lt;digit&gt; should be' part name; measured 2026-09-25; known&#95;issues.py sha1 655e5307; export 4aaee1f rows 83cf9966/ae9b703areports/known_issues/control_char_residue.tsv
target_shared_with_other_row12120The same picture (decoded pixels identical) is the target of another shipped row: one photograph or render under two upstream file names, cut into different puzzles.md5 of the decoded target pixels occurs on more than one shipped row; measured 2026-09-25; known&#95;issues.py sha1 655e5307; export 4aaee1f rows 83cf9966/ae9b703areports/known_issues/target_shared_with_other_row.tsv
validation_target_near_copy_of_train101This validation row's target is a near-copy of a training row's target (the same scene shot from a slightly shifted viewpoint), so it is not fully held out.validation target within 10 bits (of 256) of a training target by a 17x16 greyscale difference hash; the next-nearest validation-train pair on r2 is 43 bits apart; measured 2026-09-25; known&#95;issues.py sha1 655e5307; export 4aaee1f rows 83cf9966/ae9b703areports/known_issues/validation_target_near_copy_of_train.tsv

To leave the listed rows out (the snippet in the loader section downloads reports/known_issues/ with the data):

python
import glob, os
root = "<root>/thinkmorph_jigsaw"
drop = {line.split("\t")[0] for f in glob.glob(os.path.join(root, "reports/known_issues/*.tsv"))
        for line in open(f) if line.strip() and not line.startswith(("#", "row_uid\t"))}
# keep a row when its row_uid (a column of train/, meta/ and preview/) is not in drop

Measured caveats

Measured on this release by the pre-publication review (2026-09-25): problems that cannot be listed row by row (a shortcut in the options, a label convention, an upstream labelling scheme) and what the review found around the lists above. Where a caveat counts listed rows ("listed as ..."), the count is the table's, read from reports/known_issues/summary.json. Its other numbers are the review's own measurements, which no file carries: they hold for exactly these rows and are not re-measured automatically. Items marked Training-signal defect are problems in what the rows teach, not only in how they are described; no row was removed for them.

  • —Training-signal defect. The target image and the answer always show the labelled arrangement (the labelled permutation is the closest re-assembly of the input's parts to the target on all 5,799 rows), but the upstream text does not always agree with them. On 90 rows (88 train, 2 validation; listed as text_states_other_arrangement), 1.6% of the rows, the plan argues for, or the read-back says it assembled, a different arrangement, often the identity order that is not among the options, and the row then answers the label. They include read-backs and plans that quote a different arrangement in the options' own wording (e.g. 775b42a13ffe2cf2, 6ed2d5f38961ddb5, validation 4869f3448140a813), read-backs that pair '(A)' with option D's text and answer D (070c5e56e50991e1, 8df62bedf96e0c99), and two-part plans that conclude the opposite arrangement in free wording (e.g. eb30e46819a3aaae). These rows train: text says X, image shows Y, answer Y.
  • —The plan quotes the chosen option's text on 481 of the 5,799 rows. On rows flagged S4.plan_verdict_moved the sentences moved to the read-back are often the plan's comparison of the options ('If I place Part 1 below Part 2 (option A), ...') rather than its verdict, so the read-back opens with that pre-image comparison and the plan may no longer say which arrangement to draw (e.g. f0c0ba3365bd5be7). listed as readback_opens_with_moved_comparison: 172 rows, 164 train and 8 validation; the list is found from the read-back's opening sentence and also holds some rows whose moved text does end in a verdict, mostly among the three- and four-part puzzles.
  • —Training-signal defect. Upstream wrote \boxed{X} in a form mangled into control bytes; the conversion stripped the control bytes (flag S0.control_chars_stripped) but left the printable remnant, so fragments such as [boxed{B}, b[boxed{A}, cboxed{B} and boxed{A} remain in the text, nearly all in the read-back (e.g. 96f5cd52dc2240f7, d40a6f95cec102f3; listed as control_char_residue: 49 rows, 49 train and 0 validation; 99f7eab94a675c55 instead has a part number cut to '02 should be ...' in its plan), and the model is trained to write them. The letter in each boxed fragment equals the row's answer.
  • —One picture can appear under two file names: 12 rows (12 train, 0 validation; listed as target_shared_with_other_row) share their target picture with another row (e.g. ebd5964bccf870f9 and 01ce85a03268e9b2; b477f54ba9491c9d and 48ccb324817e04b8 are the same puzzle with its answer worded differently). Validation row 96fb0688103dd21d (ADEtrain00018130, listed as validation_target_near_copy_of_train) shows nearly the same street photograph as the target of training row bdcb20276b29b58b (ADEtrain00018569); it is cut into a different grid, so the check on input images does not catch it.

Size

splitrowstarget image slotsdistinct target images
train5,6545,6545,648
validation145145145

A slot is one target position in one row. Upstream has a few pictures under two file names, cut into different puzzles, so a few rows share a target picture (see Known issues).

tasktrainvalidation
jigsaw_1x21,08637
jigsaw_1x31,16330
jigsaw_2x11,09816
jigsaw_2x21,15931
jigsaw_3x11,14831

Input images per row: 1. Target images per row (the images the model is trained to generate): 1. Image corpus (source_scene_corpus): procthor 3,203, ade20k 1,821, sunrgbd 775.

Row format

One row is: input image(s) and a question, then K rounds of thought → target image (the target is the source's own ground-truth image, which the model is trained to generate), then a final thought (normally a read-back of the last target; where a source's final thought is something else, or often leaves out the answer, the source note or Known issues says so) and the answer; here K is 1. In the train config:

image_list        list<binary>  inputs first, then the K target images in order
num_input_images  int64         how many of image_list are inputs
instruction_list  list<string>  one element: system prompt + question + options
output_text_list  list<string>  K+1 elements:
  [0]   <think>plan 1</think><image_start>
  [j]   <image_end><think>plan j+1</think><image_start>
  [K]   <image_end><think>read-back</think><answer>answer</answer>
row_uid           string        join key to `meta` and `preview`

Every image is a JPEG, and no input image is larger than 512 px on its long edge (measured on this release, 2026-09-25); the size each target was stored at is target_px in meta.

On every row <answer> holds the option key (answer type mcq_letter 5,799) that the model is trained to emit (a letter for mcq_letter), while meta.answer_value (the answer column of preview) holds that option's text. Map the key through the options listed in the question before comparing the two, and score model output against <answer>.

The system prompt is ThinkMorph's VLM_THINK_SYSTEM_PROMPT from its inferencer.py, verbatim (GEN_THINK_SYSTEM_PROMPT there has the same text), including its leading and trailing newline. The markers are plain strings, not tokenizer special tokens; the prompt writes </image_end> and the data writes <image_end>, exactly as the ThinkMorph-7B checkpoint was trained.

preview shows the same rows with one column per slot: input_image_i for the inputs; for each of the K = num_steps rounds, the plan thought_j and its target target_image_j; and the read-back in thought_1 on every row.

meta holds the per-row sidecar: task, scene_id and geometry_uid (the scene and geometry keys; the split key is named in the split paragraph below), trajectory_id (a camera-path or sample label, empty where the source has none), num_steps, num_input_images, answer_type, answer_value, majority_class_rate, target_image_kind, target_px, est_tokens, licence, split (train / validation, the Hub split names), supervision_kind (full_interleaved / visual_aux / visual_only) and filter_flags. majority_class_rate is the share of the task's most frequent answer_value among its training rows: it measures answer skew and is not a guessing baseline (where a task mixes question types or each row has its own options it can be far below chance); compare scores with the text-only baselines below.

Per-row license in meta: undeclared 5,799.

Flags on released rows (filter_flags in meta and preview, comma-separated):

flagrowsmeaning
S5.replay_unsupported5,799no solver re-derives this task's answer from the trace, so S5 did not replay it
S4.plan_verdict_moved2,015the plan named the option letter; those sentences were moved to the start of the read-back
S8.phash_near_but_distinct1,373a target's perceptual hash is within 6 bits of an input image's, but its pixels differ, so it is not a copy; kept
S14.sampled_qa200chosen for the S14 human spot-check (reports/s14_sample.tsv)
S0.control_chars_stripped52control characters removed from the upstream text
S4c.freeform_unadjudicated13S4c's generic check could not match the read-back's wording (e.g. 'statement (D', 'boxed{B') to the long option text; these rows are multiple choice, not free-form, and read by hand all 13 conclude the labelled option; kept

Training with a BAGEL-family loader

Every row here has one input image (num_input_images is 1), so the stock ThinkMorph UnifiedEditIterableDataset (https://github.com/ThinkMorph/ThinkMorph: image_list[0] as input, image_list[j+1] after output_text_list[j]) and the IPT release's version (which reads num_input_images) both read it as intended. Mixed with a source whose rows have more than one input image, only a loader that reads num_input_images is correct.

The stock BAGEL edit loader (ByteDance-Seed/Bagel) cannot train these rows: it never reads output_text_list and expects each instruction_list element to be a list of paraphrases.

parquet_info.json keys each training chunk as <source>/<split>/<file>, here thinkmorph_jigsaw/train/chunk_00000.parquet, with row-group counts read from the parquet footers. The loader matches a chunk only when its key equals the path it builds, os.path.join(data_dir, file), and skips a chunk with no key without a warning: a source that is alone in its group then fails with IndexError: list index out of range, and in a mixed group it adds no rows. Download into a directory named after the source, not after the repository:

python
from huggingface_hub import snapshot_download
snapshot_download("yrlyrl/spatial-mmcot-thinkmorph_jigsaw", repo_type="dataset", local_dir="<root>/thinkmorph_jigsaw",
                  allow_patterns=["train/*", "validation/*", "parquet_info.json", "reports/known_issues/*"])

Then either run from <root> with data_dir: thinkmorph_jigsaw/train and parquet_info_path: thinkmorph_jigsaw/parquet_info.json, or rebuild the index with absolute keys and use an absolute data_dir:

python
import json, os
root = "/abs/path/to/root"                      # the directory that holds thinkmorph_jigsaw/
info = json.load(open(os.path.join(root, "thinkmorph_jigsaw", "parquet_info.json")))
info = {os.path.join(root, k): v for k, v in info.items()}
json.dump(info, open(os.path.join(root, "thinkmorph_jigsaw", "parquet_info_abs.json"), "w"))
# data_dir = os.path.join(root, "thinkmorph_jigsaw", "train")   (spelled exactly so, no trailing slash)
# parquet_info_path = os.path.join(root, "thinkmorph_jigsaw", "parquet_info_abs.json")

The Hugging Face cache (.../snapshots/<hash>/train/) or a folder named spatial-mmcot-thinkmorph_jigsaw matches no key.

num_used_data counts chunk files, not rows: the loader repeats this source's file list up to that number, lists every (file, row group) pair, and deals whole row groups out, floor(R / worldsize) to each rank and floor(that / numworkers) to each DataLoader worker. The remainder is never read. This source has 1 training chunk file holding 45 row groups of up to 128 rows, so keep num_used_data large, e.g. the 128 of ThinkMorph's interleaved_reasoning.yaml (upstream's example.yaml asks for more than GPUs x workers); every row group is then read. Set to 1 and alone in its group on 8 GPUs with 4 workers, it reads only 32 of the 45 row groups. In a run that mixes sources, give each source the same multiple of its own training chunk-file count, e.g. 128 per file (128 here): the file list is repeated up to num_used_data entries, so a flat 128 for every source would read a two-file source's rows half as often as a one-file source's.

How the rows were chosen

stagerows
upstream rows read6,000
refused before conversion (S0raw; each reason is in the table below)133
removed as benchmark evaluation items (S11)2
dropped at S5 by the phrase check ('does not make sense', 'discrepancy'); read by hand, 8 of these plans reject the arrangement they then choose or conclude a different option from the label, the others use the phrase to reject a wrong arrangement36
after conversion and per-row filters5,829
removed by S10 (none)0
removed by answer-prior balancing (S13)30
released5,799

Every removed row has one line, with its reason, in reports/:

filestepreason (the line's `flag`, or the field shown)rows
build/dropped.jsonlS0rawS0.readback_names_other_option118
build/dropped.jsonlS0rawS0.plan_names_other_option14
build/dropped.jsonlS0rawS0.plan_enumerates_options1
build/dropped.jsonlS11S11.bench_item2
build/dropped.jsonlS5S5.self_contradiction36
s13_dropped.jsonlS13step: letter24
s13_dropped.jsonlS13step: answer6

Every line of s13_dropped.jsonl has reason: prior_downsample; step names the balancing pass that removed it, and split is written train or val (the Hub's validation).

S0raw lines in build/dropped.jsonl were refused before a release row existed, so their row_uid field holds the converter's key for the upstream record, not a 16-hex row_uid; lines from later steps carry the row_uid the row had. No removed row appears in meta or preview.

S11.bench_item marks upstream rows with an image that is an evaluation item of a benchmark (2 rows: puzzles whose target, the assembled photograph, is an image of the What's Up benchmark (Kamath et al., 2023), which is not among the benchmarks we report but is kept out of training; upstream ThinkMorph Jigsaw_Assembly still contains them); they were removed during conversion, so no such item ships.

<details><summary>Per-step counters of the conversion</summary>

6,000 upstream rows were read; S0raw refused 133 before a row existed and passed 5,867 to the first step. S0 runs once more, last, on the final bytes. The reason for every refused, dropped or quarantined row is in the files above.

stepinoutdroppedquarantinedrejectedrepaired
S0raw (refused before conversion)6,0005,867001330
S115,8675,8650020
S45,8655,8650000
S4c5,8655,8650000
S55,8655,82936000
S85,8295,8290000
S95,8295,8290000
S0 (final structural check, after S9)5,8295,82900053

</details>

The train/validation split keeps rows sharing a scene_id in meta on one side, and the assignment is frozen (splits/ in the summary repository). scene_id is the image file name. Nothing links SAT renders of one ProcTHOR house, or photographs of one place, so such images can sit on both sides, and one picture can appear under two file names (examples under Known issues). S12 saw 5,829 rows under 5,829 keys, one row per key, so the split is in effect per row. No validation input image has the content of a training input image, and none is a pixel-level near-copy of one. S12 does not record per source whether that test ran, but it skips it only for a source whose spec sets split_leak_pixels: false, and no spec does; over all sources it compared 21,661 candidate pairs (perceptual hash within 6 bits) pixel by pixel and found no near-copy (checked 2026-09-25).

Answer-prior balancing (S13)

Each (task, split) group is checked separately. An answer is the answer value compared as lower-cased text without a trailing full stop, with 'farther' read as 'further' and 'nearer' as 'closer' (for multiple choice, the option text, not the letter; where the candidates are drawn in the image, as in zebrajigsaw and zebratetris, the answer is the letter itself). An answer is real when it holds at least 5 rows and 2% of the group; k is the number of real answers. Answer step: the target is max(30%, 1/k) when k >= 2, and max(30%, 1/d) over the d distinct answers when k = 1; a validation group uses the larger of its own target and its task's train target. A group is cut only when k >= 1 and its most common answer holds more than the target plus 5 percentage points; every answer is then capped at one common count, chosen so that none exceeds the target, and smaller answers keep all their rows. At the answer step, a group at or below that trigger, or with no real answer (k = 0), is left as it is, so its most common answer can hold up to the target plus 5 percentage points. A task whose train group has exactly two real answers is instead cut, in every split, so that its two largest answers have equal counts, with no trigger. Rank and label steps: then, in a group where every option value of every row is a number, the rank of the correct option among the sorted values, and after it, in a group where every trained answer is an option label, the label, are each capped by the same cut-and-trigger rule on their own counts (own target, validation included): capped, never evened out, so two labels are cut only when one exceeds 55%, and then only down to 50%. These steps can also cut groups the answer step left whole, including k = 0 groups, and can raise an answer's final share above its target; the run fails if a real answer ends above the target plus 5 percentage points. A train group of at least 20 rows in which one answer holds 90% or more fails the run. PET (exactcellspet) instead cuts each (question type x turn direction) cell to equal counts of its two answers; a PET cell that shows only one answer is removed.

tasksplitpassrulerows in → outreal answers ktargetlargest share, before → aftercut
jigsaw_1x2trainanswercap30[canon]1,086 → 1,086430.0%26.6% → 26.6%no
jigsaw_1x2trainlettercap30[letter]1,086 → 1,086250.0%51.6% → 51.6%no
jigsaw_1x2validationanswercap30[canon]41 → 37430.0%36.6% → 29.7%yes
jigsaw_1x2validationlettercap30[letter]37 → 37250.0%51.4% → 51.4%no
jigsaw_1x3trainanswercap30[canon]1,163 → 1,163630.0%16.9% → 16.9%no
jigsaw_1x3trainlettercap30[letter]1,163 → 1,163430.0%25.2% → 25.2%no
jigsaw_1x3validationanswercap30[canon]30 → 30333.3%26.7% → 26.7%no
jigsaw_1x3validationlettercap30[letter]30 → 30430.0%30.0% → 30.0%no
jigsaw_2x1trainanswercap30[canon]1,098 → 1,098430.0%26.3% → 26.3%no
jigsaw_2x1trainlettercap30[letter]1,098 → 1,098250.0%51.2% → 51.2%no
jigsaw_2x1validationanswercap30[canon]26 → 24333.3%38.5% → 33.3%yes
jigsaw_2x1validationlettercap30[letter]24 → 16250.0%66.7% → 50.0%yes
jigsaw_2x2trainanswercap30[canon]1,159 → 1,1592430.0%4.3% → 4.3%no
jigsaw_2x2trainlettercap30[letter]1,159 → 1,159430.0%25.3% → 25.3%no
jigsaw_2x2validationanswercap30[canon]37 → 37030.0%10.8% → 10.8%no
jigsaw_2x2validationlettercap30[letter]37 → 31430.0%40.5% → 29.0%yes
jigsaw_3x1trainanswercap30[canon]1,148 → 1,148630.0%17.0% → 17.0%no
jigsaw_3x1trainlettercap30[letter]1,148 → 1,148430.0%25.4% → 25.4%no
jigsaw_3x1validationanswercap30[canon]41 → 41630.0%22.0% → 22.0%no
jigsaw_3x1validationlettercap30[letter]41 → 31430.0%36.6% → 29.0%yes

S13 removed 30 rows from this source.

Text-only baselines

Accuracy of guessers that never see an image. For each task the released training rows are split into two fixed halves by a hash of row_uid; each guesser is fitted on one half and scored once on the other (one held-out half, not cross-validation; eval rows below). The reference is chance (the mean of 1 / number of options) where every row is multiple choice, and otherwise the eval-half accuracy of always giving the answer most common in the fit half (when a task's top answers are nearly tied, this need not be the task's most common answer; the line after the table gives that answer's validation score). Accuracies are recounted from the stored rates and eval rows, so they are exact. A task is flagged when a text-only guesser beats its reference by more than 0.15 (for a free-form task, a guesser other than the most common answer). A flagged task can be partly answered from the text alone; an unflagged task passed only these probes, which do not prove the text carries no answer. Report scores on every task next to this baseline.

Guessers: keywords: the most common answer per set of spatial words in the question; last_mentioned: the option named last in the question body; letter_prior: the most common answer letter; majority: the answer most common in the fit half; option_prior: the option text that won most often when shown; template: the most common answer per question wording (numbers masked, object names kept).

taskbest text-only guesseraccuracyreferencemargineval rowsflagged
jigsaw_1x2option_prior0.4950.500 (chance)-0.005547no
jigsaw_1x3letter_prior0.2430.250 (chance)-0.007575no
jigsaw_2x1option_prior0.5230.500 (chance)+0.023549no
jigsaw_2x2letter_prior0.2420.250 (chance)-0.008590no
jigsaw_3x1option_prior0.2410.250 (chance)-0.009553no

Spot-check (S14)

Pending. The S14 rows are chosen and flagged S14.sampled_qa in meta and preview; the human pass over them has not been signed off yet.

Citation

Please cite ThinkMorph (arXiv:2510.27492), which built the puzzles and their reasoning traces:

bibtex
@article{gu2025thinkmorph,
  title={ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning},
  author={Gu, Jiawei and Hao, Yunzhuo and Wang, Huichen Will and Li, Linjie and Shieh, Michael Qizhe and Choi, Yejin and Krishna, Ranjay and Cheng, Yu},
  journal={arXiv preprint arXiv:2510.27492},
  year={2025}
}

and the sources of the pictures:

  • —SAT renders (source_scene_corpus = procthor): A. Ray et al., "SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models", arXiv:2412.07755; and M. Deitke et al., "ProcTHOR: Large-Scale Embodied AI Using Procedural Generation", NeurIPS 2022.
  • —ADE20K photographs (ade20k), as ADE20K asks: B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso and A. Torralba, "Scene Parsing through ADE20K Dataset", CVPR 2017; and B. Zhou, H. Zhao, X. Puig, T. Xiao, S. Fidler, A. Barriuso and A. Torralba, "Semantic Understanding of Scenes through the ADE20K Dataset", IJCV (https://groups.csail.mit.edu/vision/datasets/ADE20K/).
  • —SUN RGB-D photographs (sunrgbd): S. Song, S. Lichtenberg and J. Xiao, "SUN RGB-D: A RGB-D Scene Understanding Benchmark Suite", CVPR 2015; and, as the SUN RGB-D page requires of every user (https://rgbd.cs.princeton.edu/), the three datasets it contains: N. Silberman, D. Hoiem, P. Kohli and R. Fergus, "Indoor Segmentation and Support Inference from RGBD Images", ECCV 2012; A. Janoch, S. Karayev, Y. Jia, J. T. Barron, M. Fritz, K. Saenko and T. Darrell, "A Category-Level 3-D Object Dataset: Putting the Kinect to Work", ICCV Workshop on Consumer Depth Cameras for Computer Vision 2011; J. Xiao, A. Owens and A. Torralba, "SUN3D: A Database of Big Spaces Reconstructed using SfM and Object Labels", ICCV 2013.

Provenance

The release files were written by our conversion code (the code repository is not public yet), scripts/convert/export.py at commit 4aaee1f4e946, from build thinkmorph_jigsaw_r2. The build was made by scripts/convert/run_source.py from the same repository at commit 4aaee1f4e946. S10, S12 and S13 ran before the export; reports/export_manifest.json pins every input the export read by SHA-1 (build_manifest_sha1, s10_keep_sha1, s12_assignments_sha1, s13_balanced_keep_sha1).

Every row removed between upstream and this release has one line, with its reason, in reports/: build/dropped.jsonl (rows refused before conversion or dropped by a conversion step, S11 benchmark items included); build/quarantine.jsonl (rows set aside by S4c because an automatic check could not match the read-back's conclusion to the label); s10_dropped.jsonl (duplicates removed by S10); s10_label_conflicts.jsonl (rows S10 withheld because another row asks the identical question, options in the same order, of the same images with a different answer); s13_dropped.jsonl (rows removed by answer-prior balancing). known_issues/ lists rows with a measured problem (see Known issues); reports/ also holds the build manifest (absolute paths cut to basenames) and counters, the S14 sample list (s14_sample.tsv: rowuid, task, split) and `exportmanifest.json. In reports/build/manifest.json, spec.sourcescenecorpus (procthor) is only the converter's fallback for rows that carry no corpus of their own; it does not describe every row. Each row's corpus is sourcescenecorpus in meta (procthor 3,203, ade20k 1,821, sunrgbd 775). Part of [yrlyrl/spatial-mmcot`](https://huggingface.co/datasets/yrlyrl/spatial-mmcot).