CoolFace
Modelpublic

prithivMLmods/SpatialBlock-7B-direct-GGUF

sourceHugging Faceapache-2.0updated 11d agoView on Hugging Face
1likes336downloads
Model Card

SpatialBlock-7B-direct-GGUF

SpatialBlock-7B-direct is an open-source multimodal model released on Hugging Face by rsoohyun under the Apache-2.0 license, developed to enhance spatial intelligence in Large Vision-Language Models (LVLMs). Based on the Qwen/Qwen2.5-VL-7B-Instruct base architecture and supported by the Hugging Face transformers library via the image-text-to-text pipeline, this checkpoint is fine-tuned on the synthetic SpatialBlock-15k dataset as presented in the paper SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem. It specializes in directly predicting solutions to intricate spatial tasks—including 3D-to-2D projection, viewpoint transformation, and structural combination—with complete training methodology, evaluation metrics, and the companion “reason” model accessible via its GitHub repository.

Model Files

File NameQuant TypeFile SizeFile Link
SpatialBlock-7B-direct.BF16.ggufBF1615.2 GBDownload
SpatialBlock-7B-direct.Q4KM.ggufQ4KM4.68 GBDownload
SpatialBlock-7B-direct.Q5KM.ggufQ5KM5.44 GBDownload
SpatialBlock-7B-direct.mmproj-bf16.ggufmmproj-bf161.36 GBDownload

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp