CoolFace
Modelpublic

gbpeck/Qwen3-0.6B-litertlm-jinja

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes56downloads
Model Card

Qwen3-0.6B — LiteRT-LM (.litertlm), Jinja chat template preserved

A LiteRT-LM (.litertlm) conversion of Qwen/Qwen3-0.6B, built with use_jinja_template=True so the model's full Jinja chat template is preserved in the artifact metadata (llm_metadata.jinja_prompt_template).

Why this exists

The stock `litert-community/Qwen3-0.6B` .litertlm was converted with the chat template stripped to a simple prefix/suffix form. On device that means:

  • —`enable_thinking` is silently inert. The thinking on/off control (passed through LiteRT-LM's extraContext) has no template conditional to bind to, so enable_thinking=false does not suppress the model's <think> reasoning.
  • —Tool calls come back as plain text. Without the template's tool-rendering block, the model never learns the <tool_call> output convention.

This conversion keeps the real template, so enable_thinking=false genuinely suppresses reasoning (verified on device: ~10× fewer completion tokens vs. on) and tool calls surface as structured tool_call events.

Details

  • —Base model: Qwen/Qwen3-0.6B (Apache-2.0)
  • —Quantization: dynamic INT8 weights (dynamic_wi8_afp32); fp32 KV cache
  • —Context: cache_length=8192 (the KV cache holds up to 8192 tokens of prompt + output; ~1.75 GiB at fp32)
  • —Backends: CPU / GPU — portable build, no NPU/vendor-specific variant
  • —Conversion: litert-torch export_hf --use_jinja_template=True --cache_length=8192
  • —File: Qwen3-0.6B.litertlm (~585 MiB)

Produced with `litert-torch`'s export_hf (CPU-only, no GPU).

License

Apache-2.0, inherited from the base model. This is a format conversion only — all weights and behavior derive from Qwen/Qwen3-0.6B; upstream attribution and NOTICE are retained.