gbpeck/Qwen3-0.6B-litertlm-jinja
Qwen3-0.6B — LiteRT-LM (.litertlm), Jinja chat template preserved
A LiteRT-LM (.litertlm) conversion of Qwen/Qwen3-0.6B, built with use_jinja_template=True so the model's full Jinja chat template is preserved in the artifact metadata (llm_metadata.jinja_prompt_template).
Why this exists
The stock `litert-community/Qwen3-0.6B` .litertlm was converted with the chat template stripped to a simple prefix/suffix form. On device that means:
- `enable_thinking` is silently inert. The thinking on/off control (passed through LiteRT-LM's
extraContext) has no template conditional to bind to, soenable_thinking=falsedoes not suppress the model's<think>reasoning. - Tool calls come back as plain text. Without the template's tool-rendering block, the model never learns the
<tool_call>output convention.
This conversion keeps the real template, so enable_thinking=false genuinely suppresses reasoning (verified on device: ~10× fewer completion tokens vs. on) and tool calls surface as structured tool_call events.
Details
- Base model: Qwen/Qwen3-0.6B (Apache-2.0)
- Quantization: dynamic INT8 weights (
dynamic_wi8_afp32); fp32 KV cache - Context:
cache_length=8192(the KV cache holds up to 8192 tokens of prompt + output; ~1.75 GiB at fp32) - Backends: CPU / GPU — portable build, no NPU/vendor-specific variant
- Conversion:
litert-torch export_hf --use_jinja_template=True --cache_length=8192 - File:
Qwen3-0.6B.litertlm(~585 MiB)
Produced with `litert-torch`'s export_hf (CPU-only, no GPU).
License
Apache-2.0, inherited from the base model. This is a format conversion only — all weights and behavior derive from Qwen/Qwen3-0.6B; upstream attribution and NOTICE are retained.
