mlx-community/gemma-4-26B-A4B-it-OptiQ-4bit
Sync chat template from google/gemma-4-31b-it (Google canonical, published 2026-07-09)
index: register the bf16 vision tower so stock VLM loaders (mlx-vlm, oMLX, LM Studio) can find it
Correct BFCL + Capability: the AST checker could not match array-typed arguments
Move OptiQ sidecars under optiq/ so *.safetensors loaders (mlx-vlm, LM Studio) skip them
fix: audio_tower conv weights to MLX channel-last layout (full-VLM loader compat)
docs: OptiQ brand spelling
card: remove em-dashes, ensure funnel
card: add mlx-optiq funnel banner + CTA
Restore vision_config + optiq_vision marker for image input
Add OptIQ vision sidecar (image+text support, v0.2.0)
Add kv_config.json (per-layer mixed-precision KV cache, 5.0 BPW target from OptiQ kv-cache sensitivity analysis). Requires optiq>=0.1.3 runtime for the RotatingQuantizedKVCache shim.
v0.1.0: 6-metric Capability Score + OptIQ vs U4 deltas
Fix chat template: emit multimodal placeholders in tool messages
v0.1.0: 79.0 Capability (no-strip + attention-aware floor + 6-domain mix)
Re-eval GSM8K with chat-template: OptIQ 94.0% / uniform 92.0% (+2.0pp)
Delete stale buggy model-00005-of-00005.safetensors (replaced by 3-shard fixed version)
Delete stale buggy model-00004-of-00005.safetensors (replaced by 3-shard fixed version)
Delete stale buggy model-00003-of-00005.safetensors (replaced by 3-shard fixed version)
Delete stale buggy model-00002-of-00005.safetensors (replaced by 3-shard fixed version)
Delete stale buggy model-00001-of-00005.safetensors (replaced by 3-shard fixed version)
Add GSM8K benchmark: OptiQ 45.5% vs uniform 4-bit 18.5% (+27.0pp)
Update Article link to https://x.com/latent_node/status/2028412948167942334
Normalize model card: remove twitter ref, lowercase 'optiq', add multimodal stripping note
Re-quantize with proper MoE expert handling: 24.5 GB → 14.5 GB, 4.94 BPW (was 8.15)
Initial upload: gemma-4-26B-A4B-it-OptiQ-4bit (mixed-precision MoE, 4.5 BPW)
initial commit
