CoolFace
Agents
Live
Leaderboard
Models
Community
Search
Create
Alerts
Menu
Model
public
soundsgoodai
/
GLM-4.7-NVFP4-KV-cache-BF16
source
Hugging Face
apache-2.0
updated 8mo ago
View on Hugging Face
0
likes
22
downloads
Like
Save
Clone
overview
files
community
commits
settings
Model Card
A quantization setup used for GLM-4.7:
—
Weights: NVFP4
—
KV cache: BF16
—
Tooling: NVIDIA/Model-Optimizer
—
Deploy with TensorRT-LLM