CoolFace
Modelpublic

prithivMLmods/SpatialBlock-4B-reason-GGUF

sourceHugging Faceapache-2.0updated 11d agoView on Hugging Face
1likes363downloads
Model Card

SpatialBlock-4B-reason-GGUF

SpatialBlock-4B-reason is an open-source multimodal model released on Hugging Face by rsoohyun under the Apache-2.0 license, developed to enhance spatial intelligence in Large Vision-Language Models (LVLMs). Based on the Qwen/Qwen3-VL-4B-Instruct base architecture and supported by the Hugging Face transformers library via the image-text-to-text pipeline, this checkpoint is fine-tuned on the synthetic SpatialBlock-15k dataset as presented in the paper SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem. It specializes in reasoning through and directly predicting solutions to intricate spatial challenges—including 3D-to-2D projection, viewpoint transformation, and structural combination—with complete training methodology, evaluation metrics, and companion models accessible via its GitHub repository.

Model Files

File NameQuant TypeFile SizeFile Link
SpatialBlock-4B-reason.BF16.ggufBF168.83 GBDownload
SpatialBlock-4B-reason.Q4KM.ggufQ4KM2.72 GBDownload
SpatialBlock-4B-reason.Q5KM.ggufQ5KM3.16 GBDownload
SpatialBlock-4B-reason.mmproj-bf16.ggufmmproj-bf16839 MBDownload

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp