CoolFace
Modelpublic

hipfire-models/qwen3-0.6b

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes
Model Card

Qwen3-0.6B for hipfire

Pre-quantized Qwen3-0.6B (LLaMA (standard attention)) for hipfire, a Rust-native LLM inference engine for AMD RDNA GPUs.

Quantized from Qwen/Qwen3-0.6B.

Files

FileQuantSizeMin VRAMSpeed (5700 XT)
qwen3-0.6b-hfq4.hfqHFQ40.4GB1GB
qwen3-0.6b-hfq4-v2.hfqHFQ4 v20.4GB1GB
qwen3-0.6b-hfq4g256.hfqHFQ4-G2560.4GB1GB

Usage

bash
# Install hipfire
curl -L https://raw.githubusercontent.com/Kaden-Schutt/hipfire/master/scripts/install.sh | bash

# Pull and run
hipfire pull qwen3:0.6b
hipfire run qwen3:0.6b "Hello"

Quantization Formats

  • HFQ4: 4-bit, 256-weight groups (0.53 B/w). Best speed.
  • HFQ6: 6-bit, 256-weight groups (0.78 B/w). Best quality. ~15% slower.

Both include embedded tokenizer and model config.

About hipfire

Rust + HIP inference engine for AMD consumer GPUs (RDNA1–RDNA4). No Python in the hot path. 9x faster than llama.cpp+ROCm on the same hardware.

License

Model weights subject to original Qwen license. hipfire engine: MIT.