CoolFace
Modelpublic

gbuzhf/Ornith-1.5-35B-A3B-TIEL-Calibrated-MTPv2-ICE-GGUF

sourceHugging Facemitupdated 16d agoView on Hugging Face
9likes5.7kdownloads
Model Card

Ornith-1.5-35B-A3B — TIEL-Calibrated MTPv2 ICE tiers

Four GGUF tiers of the original ornith-ai/Ornith-1.5-35B-A3B, with ornith-ai's trained MTPv2 head embedded, calibrated on Tiel's importance matrix, carrying the Qwen-Sharp v22.4.1 chat template.

Which one should I download?

tiersizemean KLDvs the Unsloth Dynamic ladder
`23G-ICE`22.84 GB0.0325Beats `UD-Q4_K_XL` on both axes — 14.5 % closer to bf16 and 0.37 GB smaller. Best value here.
25G-ICE24.85 GB0.0284The best file below 25 GB — nothing smaller is closer to bf16. Also beats APEX-I-Balanced (26.28 GB / 0.0345) by 12 % while being 1.4 GB smaller. Above it, UD-Q5_K_S at 25.83 GB is stronger.
21G-ICE20.85 GB0.0389Fills a 4.5 GB hole in the UD ladder. Essentially UD-Q4_K_XL quality (0.0380) at 2.36 GB less.
19G-ICE18.82 GB0.060117 % closer to bf16 than `UD-IQ4_XS` (0.0723) for +0.14 GB.

Short version: ICE wins in the 19–23 GB band and loses above 25 GB — see Method.

Measurements

KL divergence against the BF16 master this repo was built from. WikiText-2 raw test, 64 chunks, n_ctx 2048, one binary and one reference for all four files. Mean PPL(base) = 7.494953 ± 0.079236.

Sorted best → worst by overall (BF16 = 100), the same composite used on the other Ornith-1.5 cards: 0.70/(1+meanKLD) + 0.30*sameTop1.

tiersizemean KLD99% KLD99.9% KLDPPL ratiosame top-1active bpwfile bpw**overall**
25G-ICE24.84 GB0.02840.2861.1200.985193.44%7.6865.59796.1
23G-ICE22.83 GB0.03250.3351.1770.990092.85%7.5235.14395.7
21G-ICE20.84 GB0.03890.4191.5841.000192.24%7.3574.69595.1
19G-ICE18.82 GB0.06010.6162.3871.001390.33%7.1924.24093.1

Active bpw weights each tensor by how often it actually runs — routed experts at k/E — so it says where the bits went in the forward pass rather than on disk. It explains a design; it does not rank one. Ranking is on measured KLD.

Method

Every GGUF quantizer — llama.cpp's own mixes, Unsloth Dynamic, APEX — minimises the same thing for every tensor: importance-weighted error of that tensor's output, for the current token. That is correct for a tensor whose error dies with the token, and wrong for the ones whose error does not.

ICE sorts tensors by how far an error travels, then pays accordingly:

classwhat it is herewhytype
discretethe MoE routeran error flips an argmax, so a different expert runs. Not a graded loss — a categorical one.F32
recurrentSSM decay / timestep termsthe error enters a carried state and compounds along the sequenceF32
cachedattn_k, attn_vwritten to the KV cache once, re-read by every later token, never re-decidedF16
instanteverything else, incl. all 256 expertsthe error affects this token onlythe dial

The first three classes are 0.14 % of the model — 47 M parameters, 0.15 GB. Freezing them outright is a line item, not a trade-off. The recovered budget plus the whole remaining budget goes to the expert bank, uniformly across gate/up/down, with the higher type placed shallow-first.

Where the budget goes furthest. All four tiers here are Pareto-optimal against a twelve-tier comparison of the Unsloth Dynamic and APEX ladders on this model, measured on one harness against one reference: nothing published is both smaller and closer to bf16 than any of them.

The advantage is largest in the 19–23 GB band and narrows as the budget rises: above ~25 GB the expert term saturates and what remains is the dense path, which is the regime UD's "pin the dense path high" policy is built for. That is why UD-Q5_K_S and UD-Q6_K remain the strongest files on the board, and why this ladder stops at 25 GB rather than chasing them.

Full report, including seven negative results and the retraction of a rule this work itself derived and shipped: [gbuzhf/ICE-quantization](https://huggingface.co/gbuzhf/ICE-quantization)

Files and use

Ornith-1.5-35B-A3B-TIEL_Calibrated-MTPv2-{19,21,23,25}G-ICE.gguf

Each is a single file carrying the MTPv2 head (block_count=41, 753 tensors) and the Sharp template. For self-speculative decoding pass --spec-type draft-mtp; the head is embedded, no sidecar model is needed