CoolFace
Modelpublic

PrismPhi/Qwen3.5-0.8B-Radxa-Dragon-Q6A-QCS6490-QNN-NPU

sourceHugging Faceapache-2.0updated 15d agoView on Hugging Face
0likes
Model Card

Qwen3.5-0.8B for Radxa Dragon Q6A (QCS6490 QNN NPU)

日本語

Fixed-context Text/Vision model assets for the Radxa Dragon Q6A, QCS6490, HTP v68, using the standard ONNX Runtime QNN EP. Use 2K first. The optional 4K profile accepts longer input but uses more memory and can produce different answers. Both profiles use the same 1024×576, 576-token Vision encoder and tokenizer/embedding assets.

Use the companion source release v0.1.0-rc1 and its installation guide. The model directory alone is not an application: it requires the C++ host, state contracts, preprocessing and API code from that source release. These EPContext models are not loaded with transformers.from_pretrained.

Run the supplied contexts

The corresponding immutable source snapshot is 97650f6ed3a1178dcf1ff761d95e4523819bfe08. Its manifest pins model commit 05de67d2f33c0cb83e375f13788895c761d30b1e; later model-card updates do not change the asset revision used by that source.

On a working Q6A Ubuntu system with the quickstart prerequisites, obtain the source and recommended model profile:

sh
git clone --branch v0.1.0-rc1 --depth 1 https://github.com/PrismPhi/radxa-dragon-q6a-qwen3.5-0.8b-qcs6490-qnn-npu.git
cd radxa-dragon-q6a-qwen3.5-0.8b-qcs6490-qnn-npu
python3.12 -m venv .venv
.venv/bin/python -m pip install --no-deps --only-binary=:all: -r requirements-device.txt
.venv/bin/python tools/fetch_assets.py --output ./downloads --profile 2k
.venv/bin/python install.py --assets ./downloads --install ./install-2k
./install-2k/launch.sh

The full quickstart lists OS/build prerequisites, immutable downloads, hash verification, 4K switching and API examples. ORT 1.27.0 and QNN EP 2.3.0 are pinned together with actual runtime binary hashes; the compiler backend is QAIRT 2.47.0.260601114230, v68, 2 MB VTCM. Runtime wheels, SDK libraries and firmware are obtained separately and are not included in these assets.

Contents and download size

SelectionApproximate sizePurpose
Common + 2K2.053 GBRecommended runtime
Common + 4K2.055 GBOptional longer-context runtime
Common + both contexts2.903 GBSwitch between profiles without downloading common files twice
All files including frozen rebuild sources4.910 GBRecompile from the accepted quantized graphs

GB here means decimal download/storage bytes, not resident RAM. manifest.json records exact bytes and SHA-256 for every model asset. The source downloader selects only the needed profile by default.

  • —common/: original tokenizer and exact FP32 embedding lookup.
  • —models/2k/, models/4k/: Decode and full-depth Prefill wrappers around one shared binary per profile.
  • —models/vision/: patch, body and merger contexts shared by both profiles.
  • —contracts/2k/, contracts/4k/: rotary inputs, state encodings and cache-advance contracts.
  • —sources/common/: one external language-weight file and the 2k_*.onnx / 4k_*.onnx graphs beside it.
  • —sources/2k/, sources/4k/: profile manifests and Prefill I/O templates.
  • —sources/vision/: the three frozen quantized Vision graphs.
  • —upstream/: original reference configurations and chat template.

The source release documents how to compile and assemble these frozen graphs. Complete recalibration from the original checkpoint is a separate experiment; private historical calibration material is not required to execute or recompile the supplied QDQ graphs and is not redistributed.

Origin and modifications

Upstream: Qwen/Qwen3.5-0.8B, revision `2fc06364715b967f1860aea9cf38778875588b17`. Original safetensors SHA-256: `04b1c301231dd422b8860db31311ab2721511346a32cb1e079c4c4e5f1fe4696`.

The port includes fixed-shape export, accepted channel/layer quantization, stable recurrent operators, corrected logit range, device-measured Vision bias corrections, calibrated all-layer Chunk16 Prefill, exact state boundaries, fixed-context extension and grouped-query layout changes. The host supplies a bounded sequential suffix, byte-checked prefix reuse and exact candidate selection. These are modified model assets and an associated runtime policy, not an unmodified official Qwen distribution.

All learned language and Vision network operators in the QNN sessions execute on HTP with CPU EP fallback disabled. CPU tokenization, image preparation, embedding lookup, state conversion and sampling remain. One language model and one Vision encoder stay resident; the profiles are intended to run one at a time.

Observed performance and limitations

Recent workload-specific measurements on one Q6A were about 7.4 / 7.9 tokens/s Decode and 82 / 63 tokens/s Chunk16 Prefill for 2K / 4K. These are not guaranteed rates and exclude different work depending on the metric. The companion evaluation guide reports full definitions, paired comparisons, TTFT, CPU work, DMA and anonymous memory separately.

Earlier selected image/OCR cases improved from 23/30 to 28/30, with six improvements and one regression. They were diagnostic cases used during development, not a general held-out benchmark. The fresh public fixture still reads TOTAL 42 as 10, invents an OBS logo in image explanation and answers one reading task incorrectly. Both public profiles match their previous adopted installation on all eight publication-check requests, including those failures. Natural EOS is not proof of a correct answer; dense OCR, image interpretation and Japanese phrasing remain unreliable on some inputs.

Do not infer general 2K/4K equivalence from a small comparison or call the device/model's accuracy limit proven. The companion repository includes full public fixture responses and actual token IDs, the 577-case sampler differential test, an architecture diagram, bilingual porting lessons and the failure history.

License and attribution

Model derivatives retain the upstream Apache-2.0 terms; see LICENSE and NOTICE. Runtime/SDK/firmware components have separate terms and are not included. This model license does not grant rights to those external dependencies. No private user images, conversations or credentials are distributed. This is an independent, unofficial port without endorsement by Qwen, Qualcomm, Radxa or Microsoft.