CoolFace
Modelpublic

GatekeeperZA/Qwen3-VL-4B-Instruct-RKLLM-v1.2.3

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
1likes95downloads
Model Card

Qwen3-VL-4B-Instruct — RKLLM v1.2.3 (w8a8, RK3588)

RKLLM/RKNN conversion of Qwen/Qwen3-VL-4B-Instruct for Rockchip RK3588 NPU inference.

Converted with RKLLM Toolkit v1.2.3 (language model) and RKNN Toolkit (vision encoder). This is a multimodal vision-language model — it accepts both images and text as input.

Key Details

PropertyValue
Base ModelQwen/Qwen3-VL-4B-Instruct
Toolkit VersionRKLLM Toolkit v1.2.3 / RKNN Toolkit
Runtime VersionRKLLM Runtime ≥ v1.2.1 + RKNN Runtime
Quantizationw8a8 (8-bit weights, 8-bit activations)
Target PlatformRK3588
NPU Cores3
Thinking Mode❌ Disabled
Model TypeVision-Language (VLM)
LanguagesEnglish, Chinese (multilingual)

Why This Model?

Qwen3-VL-4B-Instruct is Alibaba's 4B vision-language model. It handles image understanding, visual QA, document analysis, and chart reading with strong multilingual support. Running on the RK3588 NPU enables fully local, GPU-free multimodal inference.

Compared to the smaller Qwen3-VL-2B, the 4B variant offers meaningfully better image understanding and text extraction.

Hardware Tested

  • —Orange Pi 5 Plus — RK3588, 16GB RAM, Armbian Linux
  • —RKNPU driver 0.9.8
  • —RKLLM Runtime v1.2.3

Usage

With the RKLLM API Server (VLM mode)

bash
mkdir -p ~/models/qwen3-vl-4b
cd ~/models/qwen3-vl-4b
git lfs install && git clone https://huggingface.co/GatekeeperZA/Qwen3-VL-4B-Instruct-RKLLM-v1.2.3 .

Use with GatekeeperZA/RKLLM-API-Server — the server loads both the .rkllm and .rknn files automatically when placed in the same directory.

File Listing

FileDescription
qwen3-vl-4b-instruct_w8a8_rk3588.rkllmLanguage model weights for RK3588 NPU
qwen3-vl-4b-vision_rk3588.rknnVision encoder for RK3588 NPU

Compatibility Notes

  • —Minimum runtime: RKLLM Runtime v1.2.1 + RKNN Runtime v2.x. v1.2.3 recommended.
  • —RKNPU driver: ≥ 0.9.6
  • —SoCs: RK3588 / RK3588S. Not compatible with RK3576 without reconversion.
  • —RAM: ~5.5GB loaded. Requires 8GB+ board (16GB recommended).

Acknowledgements

  • —Alibaba Qwen Team for Qwen3-VL
  • —Rockchip / airockchip for the RKLLM and RKNN toolkits
  • —Converted by GatekeeperZA