berryber09/10Eros-Max-h3-turbo-hybrid-beta4-w4a8
88.5k
MiniMax-H3 10Eros-Max beta4 (TURBO hybrid) — W4A8
W4A8 quantization of `10Eros_Max_h3_TURBO-hybrid_beta4` (bf16 source: TenStrip/10Eros-Max) — the turbo hybrid build (TenStrip/10Eros-Max).
- Format:
asym_w4a8_int8,group_size=16,convrot_groupsize=256, per-tensor Lloyd-Max codebook + fp8 group scales (Kijai / comfy-kitchen W4A8 layout). - Size: ~11.68 GB — 68.8% smaller than the ~37.5 GB bf16 source; fits a 24 GB card with headroom.
- Layers: 200 2D-linear layers quantized (96% of policy-targeted bytes); norms / first / last kept higher-precision (334 kept).
- Quantizer: comfyui-mixed-quantizer
--format w4a8 --group-size 16 --codebook-mode fit. - Verified: loads via UNETLoader (convrot) and generates a coherent image in ComfyUI before publish.
Requirements
- ComfyUI ≥ v0.31.0 (native W4A8 loader) or the
comfyui_w4a8_loader.patch. - comfy-kitchen with
AsymW4A8Int8Layout(PR #90). - CUDA SM ≥ 8.0.
Community experiment; inherits the source model's license.
