biMEMO/Qwen3.8-27B-int4-AutoRound
Update README.md
Remove internal measurement/network detail from benchmark methodology note
Update README.md
Widen benchmark table with descriptive headers and full unit names (real markdown padding does not affect HF rendering)
Widen and center-align benchmark table columns for readability
Move context-limited note to footnote for table readability
Correct benchmark numbers: measure on-host to eliminate client network overhead, fix 1-GPU gpu-memory-utilization for vLLM 0.27.1, reorder Quick Start 1/2/4 GPU
Update README.md
Update README.md
Add biMEMO Discord link
Update README.md
Add per-GPU-count Quick Start commands (1/2/4 GPUs), all benchmarked configs
Add MTP quantization approach section with A/B comparison vs full-BF16 MTP block
Upload folder using huggingface_hub
initial commit
