CoolFace
Datasetpublic

cvis-tmu/Spatial-SSRL-81k

Spatial-SSRL-81k 📖Paper| 🏠Github |🤗Spatial-SSRL-7B Model | 🤗Spatial-SSRL-3B Model | 🤗Spatial-SSRL-Qwen3VL-4B Model | 🤗Spatial-SSRL-81k Dataset | 📰Daily Paper Spatial-SSRL-81k is a training dataset for enhancing spatial understanding in large vision-language models. It contains 81,053 samples of five pretext tasks for self-supervised learning, offering simple, intrinsic supervision that scales RLVR efficiently. 📢 News 🚀 [2026/04/05] We have released… See the full description on the dataset page: https://huggingface.co/datasets/cvis-tmu/Spatial-SSRL-81k.

sourceHugging Facemitupdated 11d agoView on Hugging Face
0likes174downloads
SpatialSSRL_eval16.json39 linesDownload Raw Back to root
1[2  {3    "instruction": "Your output must be formatted as '<think>thought process as you decide on an answer</think><answer>final answer</answer>. Start your response with the '<think>' tag and then write your thought process as you work to answer the question based on the image contents. Once you are confident that you have the correct answer, close your thought process with </think>, open your answer with <answer>, copy in only your final answer, then close your answer (and your output in general) with '</answer>'.",4    "input": "The original image <image> has three regions marked as 1, 2, and 3. Consider the content, positional relationships, depths of the three regions and other cues, and sort the depths of the three regions from smallest to largest from the camera, where the format of the answer is 'Smallest-Middle-Largest'.",5    "output": "<think>To determine the depths of the regions marked as 1, 2, and 3, I'll analyze their positions and apparent distances based on cues such as overlaps and known object sizes.\n\n1. **Region 1**: This is a light fixture hanging from the ceiling. It's likely closer to the camera compared to other ceiling elements because light fixtures typically hang down from the ceiling.\n\n2. **Region 2**: This area is part of the ceiling tiles. Since it's on the ceiling itself and does not protrude further downwards like the light fixture, it should be further than the light fixture but closer than the wall or anything attached to the wall.\n\n3. **Region 3**: This appears to be part of the upper wall or ceiling grid, possibly an air vent or similar nearer the wall. Given the perspective, it should be the furthest from the camera among the three, as walls typically recede further back relative to ceiling fixtures.\n\nFrom this analysis, the sorting from closest to farthest from the camera is:\n\n\\[\n</think><answer>1-2-3</answer>",6    "images": [7      "coldstart_SFT_images/img_18.jpg"8    ],9    "question_id": "SpatialSSRL_coldstart_18"10  },11  {12    "instruction": "Your output must be formatted as '<think>thought process as you decide on an answer</think><answer>final answer</answer>. Start your response with the '<think>' tag and then write your thought process as you work to answer the question based on the image contents. Once you are confident that you have the correct answer, close your thought process with </think>, open your answer with <answer>, copy in only your final answer, then close your answer (and your output in general) with '</answer>'.",13    "input": "The image has been divided into 4 labeled 0, 1, 2, and 3. One of the four patches is flipped either horizontally or vertically. The patches are given in <image>. Based on visual clues such as continuity, alignment, and context, answer the correct patch that is flipped and the direction the flip, where the format is 'Label-Direction'. The direction can only be 0(flipped vertically) or 1(flipped horizontally). ",14    "output": "<think>Let's analyze the image patches:\n\n- **Patch 0**: This section has a clear sky.\n- **Patch 1**: This section also has a clear sky, and the airplane is visible.\n- **Patch 2**: This patch shows the lower section of the scene with houses.\n- **Patch 3**: This patch shows the airplane as well.\n\nNow, consider the alignment and continuity:\n\n- The airplane in patches 1 and 3 seems to be aligned.\n- However, patch 2 has part of the airplane's nose pointing left, which seems incorrect given the rest of the image context.\n\nNext, identify if there is a flip:\n\n- Flipping patch 2 horizontally would make the nose of the airplane align naturally pointing right, matching the rest of the airplane visible in patches 1 and 3.\n\nThus, the flipped patch is patch 2, flipped horizontally.\n\n\\(</think><answer>2-1</answer>",15    "images": [16      "coldstart_SFT_images/img_24.jpg"17    ],18    "question_id": "SpatialSSRL_coldstart_24"19  },20  {21    "instruction": "Your output must be formatted as '<think>thought process as you decide on an answer</think><answer>final answer</answer>. Start your response with the '<think>' tag and then write your thought process as you work to answer the question based on the image contents. Once you are confident that you have the correct answer, close your thought process with </think>, open your answer with <answer>, copy in only your final answer, then close your answer (and your output in general) with '</answer>'.",22    "input": "I've taken an image <image> and there are two regions marked as 1, and 2 on the image. Assume that there is a camera at position '2' and it's facing to the left of the image. According to the camera, where is the region marker '1'? A. Front B. Left C. Right D. Back. Consider cues such as depth, orientation, and 3D spatial relationship. The final answer should be chosen from 'A', 'B', 'C', and 'D'.",23    "output": "<think>To determine the relative position of region marker '1' as seen from the camera at position '2' facing to the left, let's consider the orientation and spatial relationships:\n\n1. **Camera Orientation**: The camera at position '2' is facing to the left of the image.\n\n2. **Spatial Arrangement**: The regions '1' and '2' appear to be on the same vertical plane of the structure, but region '2' is further to the right compared to region '1'.\n\n3. **Perspective View**: Since the camera is facing left, anything directly to its left will appear as 'front' or 'left' to it from its perspective.\n\nIn this setup, region '1' is relatively to the left of region '2'. Given the camera faces left, region '1' would appear towards the camera's front due to the small lateral distance and orientation. Hence, the correct answer is 'A. Front'.\n\n\\(</think><answer>A. Front</answer>",24    "images": [25      "coldstart_SFT_images/img_25.jpg"26    ],27    "question_id": "SpatialSSRL_coldstart_25"28  },29  {30    "instruction": "Your output must be formatted as '<think>thought process as you decide on an answer</think><answer>final answer</answer>. Start your response with the '<think>' tag and then write your thought process as you work to answer the question based on the image contents. Once you are confident that you have the correct answer, close your thought process with </think>, open your answer with <answer>, copy in only your final answer, then close your answer (and your output in general) with '</answer>'.",31    "input": "The image has been divided into 3 shuffled vertical strips labeled 0, 1, and 2. The shuffled patches are given in <image>. Based on visual clues such as continuity, alignment, and context, answer the correct arrangement of the strips to restore the original image, where the format is 'Top-Middle-Bottom'. ",32    "output": "<think>To solve the arrangement of the strips:\n\n1. **Visual Clues**:\n   - Strip 0 features a platform with a person walking and part of a train.\n   - Strip 1 contains the station platform with multiple people and a view of the station building on the left.\n   - Strip 2 has a clear sky and a structure that connects with a roof seen in Strip 1.\n\n2. **Analyzing Connections**:\n   - The platform seen in Strip 0 can logically continue into the busier platform seen in Strip 1, indicating that Strip 0 should come after Strip 1 in the sequence.\n   - Strip 2 shows the top of a building that could align with the rooftop visible in Strip 1, suggesting that Strip 2 should be on top.\n\n3. **Logical Order**:\n   - Based on how rooftops and platforms align, Strip 2 should be at the top to show the entire structure of the station building.\n   - Consequently, Strip 1 should come in the middle, continuing the scene with more of the station visible.\n   - Lastly, Strip 0 completes the sequence at the bottom, with the continuation of the platform and the train.\n\nThus, the correct order is: \\(</think><answer>2-1-0</answer>",33    "images": [34      "coldstart_SFT_images/img_39.jpg"35    ],36    "question_id": "SpatialSSRL_coldstart_39"37  }38]39