cloudyu/hy3-gguf
hy3-gguf — GGUF weights for Tencent Hy3 (tencent/Hy3)
GGUF-format weights for tencent/Hy3 (HYV3ForCausalLM, model_type: hy_v3), a 295B-parameter / 21B-active-parameter Mixture-of-Experts model from Tencent's Hunyuan ("Hy") team (`tencent/Hy3`).
These files were produced by the [`hy3`](https://github.com/yuhai-china/hy3) converter (hy3-convert) and are meant to be run with the `hy3` inference engine, a from-scratch C/Metal/CUDA implementation.
## ⚠️ This GGUF does NOT work with llama.cpp Despite the.ggufextension, these files are only usable by the [`hy3`](https://github.com/yuhai-china/hy3) engine.llama.cpp,ollama,LM Studio,text-generation-webui,koboldcpp, and any other llama.cpp-based tool cannot load these files. Three independent reasons: 1. Unknown architecture. The metadata declaresgeneral.architecture = "hy_v3". llama.cpp only knowshunyuan-moe,hunyuan-dense,hunyuan_vl— loading aborts withunknown model architecture: 'hy_v3'. 2. Custom metadata keys. All hyperparameters use thehy_v3.*prefix (hy_v3.block_count,hy_v3.expert_count, …), which llama.cpp does not look up. 3. Non-fused expert tensors. Experts are stored one tensor per expert (blk.N.ffn_gate_exps.0.gate_proj.weight,…1…, … — 46080 tensors), whereas llama.cpp expects experts fused into a single stacked 3D tensor per layer. This is a fundamentally different on-disk layout. This is a custom GGUF readable only by thehy3loader. Do not open issues against llama.cpp for these files.
How to run
Use the hy3 engine: <https://github.com/yuhai-china/hy3>
git clone https://github.com/yuhai-china/hy3
cd hy3
make # macOS builds the Metal backend automatically
# download a GGUF from this repo, then:
./run_metal.sh -m /path/to/hy3_q4k_mixed.gguf -p "The capital of France is" -experts 8Testing scope: thehy3engine's performance work and benchmarks were developed and verified only on macOS / Apple Silicon (Metal backend), measured on an M2 Ultra (~20–27 tok/s decode depending on-experts). The CPU and CUDA backends exist in the source but were not exercised as part of that work — treat them as untested.
Files / quantization
The mixed-precision GGUF follows this scheme (see hy3_convert.c):
Model facts
The engine supports a runtime top-k experts override (-experts 1..8) to trade quality for speed. On a small 13-question code/reasoning eval (greedy, no-think): experts=8 → 10/13, experts=4 → 7/13. Default is 8.
Chat template
Hy3 is instruction-tuned and expects the Hunyuan V3 chat format (the hy3 engine applies it automatically; use --raw to bypass). Single user turn, no-think:
<|hy_begin_of_sentence:opensource|><|reasoning_mode:opensource|>reasoning_effort:no_think<|hy_User:opensource|>{prompt}<|hy_Assistant:opensource|><think:opensource></think:opensource>Generation stops on <|hy_eos:opensource|> (120025), <|hy_endofsentence|> (120001), or <|hy_EOT|> (120008).
License & attribution
Weights derive from `tencent/Hy3`; refer to the upstream repository for the governing model license. This is an unofficial community conversion, not affiliated with or endorsed by Tencent.
