CoolFace
Modelpublic

biMEMO/Qwen3.8-27B-int4-AutoRound

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
5likes5kdownloads
15 commits on main
2d7425f1mo ago

Update README.md

MinerNinja
797e79f1mo ago

Remove internal measurement/network detail from benchmark methodology note

MinerNinja
d66a6ae1mo ago

Update README.md

MinerNinja
8d5e1381mo ago

Widen benchmark table with descriptive headers and full unit names (real markdown padding does not affect HF rendering)

MinerNinja
98dbd231mo ago

Widen and center-align benchmark table columns for readability

MinerNinja
8fd22061mo ago

Move context-limited note to footnote for table readability

MinerNinja
a651baf1mo ago

Correct benchmark numbers: measure on-host to eliminate client network overhead, fix 1-GPU gpu-memory-utilization for vLLM 0.27.1, reorder Quick Start 1/2/4 GPU

MinerNinja
17c77021mo ago

Update README.md

MinerNinja
a060d041mo ago

Update README.md

MinerNinja
de37f941mo ago

Add biMEMO Discord link

MinerNinja
93b4bb21mo ago

Update README.md

MinerNinja
2057d001mo ago

Add per-GPU-count Quick Start commands (1/2/4 GPUs), all benchmarked configs

MinerNinja
076dcbe1mo ago

Add MTP quantization approach section with A/B comparison vs full-BF16 MTP block

MinerNinja
76649791mo ago

Upload folder using huggingface_hub

MinerNinja
a875bf71mo ago

initial commit

MinerNinja