gbuzhf/Ornith-1.5-35B-A3B-TIEL-Calibrated-MTPv2-ICE-GGUF
Ornith-1.5-35B-A3B — TIEL-Calibrated MTPv2 ICE tiers
Four GGUF tiers of the original ornith-ai/Ornith-1.5-35B-A3B, with ornith-ai's trained MTPv2 head embedded, calibrated on Tiel's importance matrix, carrying the Qwen-Sharp v22.4.1 chat template.
Which one should I download?
Short version: ICE wins in the 19–23 GB band and loses above 25 GB — see Method.
Measurements
KL divergence against the BF16 master this repo was built from. WikiText-2 raw test, 64 chunks, n_ctx 2048, one binary and one reference for all four files. Mean PPL(base) = 7.494953 ± 0.079236.
Sorted best → worst by overall (BF16 = 100), the same composite used on the other Ornith-1.5 cards: 0.70/(1+meanKLD) + 0.30*sameTop1.
Active bpw weights each tensor by how often it actually runs — routed experts at k/E — so it says where the bits went in the forward pass rather than on disk. It explains a design; it does not rank one. Ranking is on measured KLD.
Method
Every GGUF quantizer — llama.cpp's own mixes, Unsloth Dynamic, APEX — minimises the same thing for every tensor: importance-weighted error of that tensor's output, for the current token. That is correct for a tensor whose error dies with the token, and wrong for the ones whose error does not.
ICE sorts tensors by how far an error travels, then pays accordingly:
The first three classes are 0.14 % of the model — 47 M parameters, 0.15 GB. Freezing them outright is a line item, not a trade-off. The recovered budget plus the whole remaining budget goes to the expert bank, uniformly across gate/up/down, with the higher type placed shallow-first.
Where the budget goes furthest. All four tiers here are Pareto-optimal against a twelve-tier comparison of the Unsloth Dynamic and APEX ladders on this model, measured on one harness against one reference: nothing published is both smaller and closer to bf16 than any of them.
The advantage is largest in the 19–23 GB band and narrows as the budget rises: above ~25 GB the expert term saturates and what remains is the dense path, which is the regime UD's "pin the dense path high" policy is built for. That is why UD-Q5_K_S and UD-Q6_K remain the strongest files on the board, and why this ladder stops at 25 GB rather than chasing them.
Full report, including seven negative results and the retraction of a rule this work itself derived and shipped: [gbuzhf/ICE-quantization](https://huggingface.co/gbuzhf/ICE-quantization)
Files and use
Ornith-1.5-35B-A3B-TIEL_Calibrated-MTPv2-{19,21,23,25}G-ICE.ggufEach is a single file carrying the MTPv2 head (block_count=41, 753 tensors) and the Sharp template. For self-speculative decoding pass --spec-type draft-mtp; the head is embedded, no sidecar model is needed
