CoolFace
Modelpublic

kingjones777/Ling-3.0-tiny-base-midtrain-ROCmFP4-COHERENT-GGUF

sourceHugging Facemitupdated 28d agoView on Hugging Face
0likes152downloads
Model Card
### ๐Ÿ”ง Runtime: build the ROCmFPX fork below Stock llama.cpp will not load this file. You need both the `bailing-hybrid` architecture and the ROCmFP4 tensor types in one tree. Upstream `charlie12345/ROCmFPX` has the ROCmFP4 types but not bailing-hybrid. Our fork has both: [`kingjones30/ROCmFPX`](https://github.com/kingjones30/ROCmFPX) โ€” a fork of charlie12345/ROCmFPX, branch main. ``bash git clone https://github.com/kingjones30/ROCmFPX.git cd ROCmFPX cmake -B build -DGGML_HIP=ON -DGPU_TARGETS=gfx1151 -DGGML_NATIVE=ON -DCMAKE_BUILD_TYPE=Release cmake --build build --target llama-server llama-quantize -j$(nproc) ` Verified 2026-08-27 on gfx1151: clean clone โ†’ **0 build errors** โ†’ llama-server loads a bailing-hybrid` ROCmFP4 GGUF from this family and generates coherent text.

Ling-3.0-tiny-base-midtrain โ€” ROCmFP4 for AMD Strix Halo (gfx1151)

4-bit ROCmFP4 quantisation of Ling-3.0-tiny-base-midtrain, one of the six Ling-3.0 base/training checkpoints inclusionAI released on 2026-08-20. The multi-token-prediction (MTP) head is preserved.

โš ๏ธ This is a base checkpoint, not an instruct model

Released for continued pretraining, domain adaptation and fine-tuning. Not instruction-tuned. It ships a chat_template.jinja, but that is a tokenizer asset and does not make the weights conversational. Prompt it as a text continuation model. For chat, use inclusionAI/Ling-3.0-tiny.

The file

ftype102 โ€” Q4_0_ROCMFP4_COHERENT
size4,769,679,040 bytes
architecturebailing-hybrid โ€” hybrid KDA linear attention + MLA
tensors549 ยท block_count 25 (24 layers + 1 MTP layer)
q_lora_rank256 โ€” low-rank compressed queries
context262,144

Head protection, verified in the finished file:

output.weight        Q6_K     1536 x 157184
token_embd.weight    Q6_K     1536 x 157184

tie_word_embeddings is false, so --output-tensor-type does real work here โ€” the COHERENT tier alone leaves output.weight at 4-bit. Both heads were forced to Q6_K and audited on exact tensor names after the build.

โš™๏ธ Requires a patched llama.cpp โ€” details

Ling-3.0-tiny sets q_lora_rank: 256, so its MLA layers use a two-stage compressed query (q_a_proj โ†’ RMS norm โ†’ q_b_proj). Ling-3.0-flash sets q_lora_rank: null and uses a single wide q_proj. A bailing-hybrid implementation written against flash therefore cannot load tiny.

The patch, against ROCmFPX, mirrors the existing DeepSeek-V2 low-rank query path:

filechange
gguf-py/gguf/tensor_mapping.pymap attention.q_a_proj / q_b_proj / q_a_layernorm โ†’ ATTN_Q_A / ATTN_Q_B / ATTN_Q_A_NORM
gguf-py/gguf/constants.pyadd those three tensors to MODEL_ARCH.BAILING_HYBRID (TensorNameMap skips anything not in the arch list)
convert_hf_to_gguf.pyemit add_q_lora_rank() when set; the null path is unchanged
src/models/bailing-hybrid.cppread Q_LORA_RANK as optional; when n_lora_q > 0 create wq_a / wq_b / attn_q_a_norm and run q_a โ†’ RMS โ†’ q_b at both graph sites, else keep the wide wq

Because the KV is read as optional, flash GGUFs (which lack it) keep n_lora_q = 0 and take the original path unchanged. ~51 lines across 5 files.

Verify a build carries it with strings libllama.so | grep bailing-hybrid โ€” the architecture table lives in the shared library, not the thin CLI binary.

Measured throughput

AMD Ryzen AI Max+ 395, Radeon 8060S (gfx1151), ROCm 7.2.4, 128 GB unified memory. llama-cli, -dio -ngl 999 -st -c 2048 -n 512 --temp 0 --seed 1234, 3 repetitions.

configflagsgeneration (median)runs
no drafter--spec-type none107.8 t/s107.9 / 107.7 / 107.8
MTP n-max 3--spec-type draft-mtp --spec-draft-ngl 999 --spec-draft-n-max 3113.3 t/s113.3 / 109.1 / 115.1

+5.1% from MTP. Enable MTP. n-max 3 is worth +5.1% on this checkpoint.

The three Ling-3.0-tiny base checkpoints do not agree on this. Measured on identical hardware with identical flags: tiny-base +7.5%, tiny-base-midtrain +5.1%, tiny-base-30T โˆ’4.4%. Every comparison was range-disjoint against a 0.2โ€“0.7% baseline spread. The MTP draft head is trained alongside the model, so its quality is a property of the training checkpoint, not of the architecture โ€” and a checkpoint can be fully competitive on quality and speed while shipping a draft head that is a net negative. Measure before enabling it.

Sample output

Continuation from "The history of mathematics begins in ancient times. One of the earliest known":

 The history of mathematics begins in ancient times. One of the earliest known mathematical texts is the Plimpton 322 tablet, which dates back to around 1800 BCE in Mesopotamia. This tablet contains a list of Pythagorean triples, which are sets of three positive integers that satisfy the Pythagorean theorem. The ta

Not measured

Perplexity is not published for this build. No perplexity figure is quoted because none was completed. Quality evidence is the coherence check above plus the tensor-level audit.

Provenance

Converted from inclusionAI/Ling-3.0-tiny-base-midtrain at revision 2c07d29ee370592bd804a0a865db5503d8a4bfa2 to BF16 GGUF (549 tensors), then quantised to ftype 102 with --output-tensor-type q6_K. Licence MIT, inherited from the base model.