irl-kit/SPARC-Qwen3.5-4B
SPARC-Qwen3.5-4B
Qwen3.5-4B fully fine-tuned for embodied spatial reasoning using VQA data generated from SPARC annotations.
Training data
The training mixture contains SPARC-generated VQA data from ours_adaptive_det_soft_snr_sp8, FSD, RoboPoint, and LLaVA-OneVision2. SPARC samples use an annotation-quality threshold of 0.97, are sorted by score, and are capped at 700 samples per object. The paper reports 1,159,047 training pairs for this mixture.
The SPARC training data is the filtered release subset, which needs no further SPARC filtering. Its filtering script and release mixture manifest are in irl-kit/SPARC-VQA-Raw. FSD, RoboPoint, LLaVA-OneVision2, and EO-1.5M remain their respective upstream datasets.
Prompting
Prompt formatting is important for these models. Use the bundled chat_template.jinja through processor.apply_chat_template(..., add_generation_prompt=True), with one user turn containing the image(s) followed by the text question. Disable thinking/reasoning mode to match evaluation.
For a single point, append exactly:
Output the point coordinates in JSON format like [{"point_2d": [x, y], "label": "target"}]. Use integer coordinates between 0 and 1000.For a trajectory or multiple points, append exactly:
Return only a JSON list like [{"point_2d": [x1, y1], "label": "point_1"}, {"point_2d": [x2, y2], "label": "point_2"}, ...]. Use integer coordinates between 0 and 1000.Training
The vision encoder is frozen and the vision projector is trainable. The model was fully fine-tuned for one epoch with a learning rate of 2e-5 and a maximum sequence length of 5600.
Evaluation
This is the paper's primary 4B model. The paper reports a 62.7 pointing/VQA average for this model. The local full benchmark evaluation associates the released weights with aggregate score 0.698.
Citation
@article{blank2026sparc,
title={SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale},
author={Blank, Nils and others},
journal={arXiv preprint arXiv:2606.13497},
year={2026}
}