CoolFace
Modelpublic

berryber09/10Eros-Max-h3-turbo-hybrid-beta4-w4a8

sourceHugging Faceotherupdated 29d agoView on Hugging Face
8likes8.5kdownloads
Model Card

MiniMax-H3 10Eros-Max beta4 (TURBO hybrid) — W4A8

W4A8 quantization of `10Eros_Max_h3_TURBO-hybrid_beta4` (bf16 source: TenStrip/10Eros-Max) — the turbo hybrid build (TenStrip/10Eros-Max).

  • —Format: asym_w4a8_int8, group_size=16, convrot_groupsize=256, per-tensor Lloyd-Max codebook + fp8 group scales (Kijai / comfy-kitchen W4A8 layout).
  • —Size: ~11.68 GB — 68.8% smaller than the ~37.5 GB bf16 source; fits a 24 GB card with headroom.
  • —Layers: 200 2D-linear layers quantized (96% of policy-targeted bytes); norms / first / last kept higher-precision (334 kept).
  • —Quantizer: comfyui-mixed-quantizer --format w4a8 --group-size 16 --codebook-mode fit.
  • —Verified: loads via UNETLoader (convrot) and generates a coherent image in ComfyUI before publish.

Requirements

  • —ComfyUI ≥ v0.31.0 (native W4A8 loader) or the comfyui_w4a8_loader.patch.
  • —comfy-kitchen with AsymW4A8Int8Layout (PR #90).
  • —CUDA SM ≥ 8.0.

Community experiment; inherits the source model's license.