litert-community/LFM2.5-1.2B-JP
Chat template: accept the 0.18 content-parts form (string form unchanged); weights, tokenizer and executor metadata byte-identical
Card: measured device / browser results (edge-compat, 2026-09)
Card: definition line (LiteRT, 2026-09)
manifest: same-device CPU control rows on the Galaxy S26 for the flagship GPU recommendation (S7 backfill 2026-09-05); the recommendation is rewritten from the measured pair (measured win, kept for prefill, or reversed to cpu)
Card: base_model_relation: quantized (list under the base model's Quantizations)
Add measured Raspberry Pi 5 (CPU) section
litertlm_manifest.json: add the _int8_gpu variant
Card: document the _int8_gpu file; correct the quantization sentence in Conversion notes
Add LFM2.5-1.2B-JP_int8_gpu.litertlm: GPU-capable int8 (same weights as _int8, re-exported on the litert-torch 0.9.3 lineage)
README: iOS Metal verified on iPhone 17 Pro — set maxNumTokens=1024 (the #3129 failure was a context-sizing issue, not a runtime bug)
Document the _int4_gpu variant: backend, measured GPU speeds, conversion notes
Add GPU-capable int4 variant (litert-torch 0.9.3 export; full OpenCL delegation)
Card: state the actual GPU-delegation blocker (INT64 ShortConv ops, 536/579 delegated)
Card: measured Performance table (M4 Max CPU/GPU + on-device where measured), accuracy note
Card: note the litert-lm 0.15 ExecutorMetadata update (weights unchanged)
Add ExecutorMetadata section required by litert-lm >= 0.15 (weights unchanged; still runs on 0.14)
Add ExecutorMetadata section required by litert-lm >= 0.15 (weights unchanged; still runs on 0.14)
LFM2.5-1.2B-JP LiteRT-LM: int8 linears-only (GSM8K 65% vs bf16 63 = parity; conv-int8 costs this tune ~9pt) + int4 736MB; CPU backend; conv-state prefill fix baked in
initial commit
