barozp/Qwen3.8-27B-Opus-Distill-v2-FP8
41.1k
Model card: cross-link MLX sibling releases
Model card: add paired BF16-vs-FP8 benchmark results + full test environment
Model card: full rewrite (changelog, native-FP8 validation matrix, sglang/MTP serving guide)
Model card: serving notes (config fix revision, sm120 flashinfer workarounds)
Fix modules_to_not_convert style: bare module paths (fused-shard matching)
Fix modules_to_not_convert: add missing linear_attn.in_proj_ba entries
Upload README.md with huggingface_hub
FP8 (block-wise e4m3, dynamic activation) quantization
initial commit
