electron-rare/Qwen3.6-35B-A3B-Q4_K_M-gguf
Qwen3.6-35B-A3B Q4KM GGUF
Q4KM quantization of `Qwen/Qwen3.6-35B-A3B` for llama.cpp inference.
- Base architecture: Qwen3.5 MoE family (35B total params, ~3B active per token)
- Quantization: Q4KM (~4.5 bits per weight, balanced quality/size)
- Use case: local inference on commodity GPU (~22 GB VRAM) or Apple Silicon Metal
Usage
llama-server -m Qwen3.6-35B-A3B-Q4_K_M.gguf --port 8000Provenance
Quantized from the original Qwen3.6-35B-A3B weights using llama.cpp/convert_hf_to_gguf.py + llama-quantize. Apache-2.0 license inherited.
Org context
This repo is part of electron-rare, the legacy research org. Production ailiance artifacts now live at `Ailiance-fr` (post-2026-05-11 carve-out).
Bench (2026-05-11)
kicad-sch generation (5-axis N3 evaluation)
Used as baseline cell ("base Qwen3.6 sans LoRA") for the kicad-sch sweep. All 5 axes zero - confirms the diagnostic finding that base models emit well-formed S-expressions (EOS 3-4/5 prompts) but fail KiCad v10 strict schema validation (missing required headers, lib_symbols).
Cross-validated by iact-bench-kicad Docker validator on electron-server: 100% agreement (5/5 FAIL Failed to load schematic).
Spec: https://github.com/ailiance/ailiance-bench/blob/main/docs/superpowers/specs/2026-05-11-kicad-sch-gap-design.md
W3 lm-eval-harness (2026-05-11)
The arc_easy result deserves investigation - peer 4-bit MLX variants score 0.80+ on the same task.
Bench comparison (2026-05-11)
kicad-sch generation (this base model used as L3 baseline)
Cross-validated by iact-bench-kicad Docker validator on electron-server: 100% agreement (5/5 FAIL).
W3 lm-eval-harness (2026-05-11)
Spec: <https://github.com/ailiance/ailiance-bench/blob/main/docs/superpowers/specs/2026-05-11-kicad-sch-gap-design.md>
