CoolFace
Modelpublic

teamblobfish/DeepSeek-V4-Pro-GGUF

sourceHugging Facemitupdated 5mo agoView on Hugging Face
15likes2.7kdownloads
16 commits on main
4c147205mo ago

README: drop multi-GPU caveats list (kept recommended config)

cchuter
5510ba25mo ago

README: add WIP multi-GPU CUDA section with recommended layer-split config

cchuter
73f32a85mo ago

README: CUDA build — add SM table, BUILD_TYPE=Release, multi-GPU needs both CXX+CUDA flags

cchuter
5366b795mo ago

README: drop #21149 gate reference, lead with fork install

cchuter
d2dbe045mo ago

README: correct -cmoe technical explanation (op_offload + peak-liveness compute buffer) per codex audit

cchuter
b111a905mo ago

README: document -cmoe + -ub 128 recommendation (fairydreaming's RTX PRO 6000 OOM report)

cchuter
5cec21d5mo ago

README: document -DGGML_SCHED_MAX_SPLIT_INPUTS=128 flag for multi-GPU CUDA builds

cchuter
cf4ced25mo ago

README: CUDA support landed (19/19 on RTX 5090), branch renamed to feat/v4-port-cuda

cchuter
abaecff5mo ago

Add Q8_0 shards

cchuter
83ab6145mo ago

Add Q4_K_M-XL shards

cchuter
19ed4b15mo ago

README: Q2_K-XL gate-tools ✓ pass; add real timing numbers

cchuter
5d5fba55mo ago

Add Q2_K-XL shards

cchuter
d8f72535mo ago

README: explicit CUDA / non-Metal backends not supported (Metal-only ops)

cchuter
8a81b3c5mo ago

Add README

cchuter
04e67405mo ago

Add DSML chat template

cchuter
6fc76a35mo ago

initial commit

cchuter