CoolFace
Modelpublic

kingjones777/PhoneLLM-Alpha-1-ROCmFP4-GGUF

sourceHugging Facebsd-2-clauseupdated 23d agoView on Hugging Face
0likes326downloads
Model Card

PhoneLLM Alpha 1 — ROCmFP4 for AMD Strix Halo (gfx1151)

The model here is not our work. PhoneLLM Alpha 1 is by [Daily](https://www.daily.co/) / the [Pipecat](https://www.pipecat.ai/) team`pipecat-ai/phonellm-alpha-1` — a full-parameter fine-tune of [NVIDIA Nemotron 3 Nano 30B-A3B](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16). This repository adds only the ROCmFP4/ROCmFPX quantisation ladder for AMD Strix Halo and the measurements below. Go star their repo. Pipecat also ship an official NVFP4 build for NVIDIA Blackwell: `pipecat-ai/phonellm-alpha-1-nvfp4`.

A hybrid Mamba-Transformer MoE voice-agent model — 30B total, 3.5B active — quantised to run on a Ryzen AI Max+ 395 (Radeon 8060S, gfx1151). One binary, both backends: HIP (ROCm) and Vulkan are a runtime `-dev` flag, not a rebuild.

PhoneLLM is built for one job: call the right tool at the right time, with thinking disabled, at phone-call latency. That shapes how we verified it — see Verification.

Which file should I use?

FileftypeSizeHeadNotes
`Q4_0_ROCMFP4_STRIX_LEAN`10615.91 GiBq8_0Flagship — start here. Smallest tier that keeps full tool behaviour. Also published on its own: `PhoneLLM-Alpha-1-ROCmFP4-STRIX-LEAN-GGUF`
Q4_0_ROCMFP4_FAST10315.83 GiBq8_0Smallest file; same tool score as the flagship
Q4_0_ROCMFP4_COHERENT10216.91 GiBq8_0Highest-precision 4-bit tier
Q6_0_ROCMFPX_AGENT11427.26 GiBq8_06-bit
Q8_0_ROCMFPX11130.37 GiBq8_0Plain Q8 reference
Q8_0_ROCMFPX_AGENT11530.84 GiBq8_08-bit, agent-tuned tensor map
Take the 15.91 GiB flagship. Across our probe the 30 GiB Q8 tiers score no better than the 16 GiB 4-bit tiers (see Verification). On a 128 GB Strix Halo that leaves real headroom to co-host your ASR and TTS models on the same box — which is the point, since PhoneLLM is the LLM stage of a voice pipeline, not a speech model (it is text-in / text-out; you still need STT and TTS).

Quick start

bash
llama-server -m PhoneLLM-Alpha-1-Q4_0_ROCMFP4_STRIX_LEAN.gguf \
  -dev ROCm0 -fa on -ngl 999 -fit off -np 1 \
  -c 32768 -b 4096 -t 8 --jinja \
  --host 0.0.0.0 --port 8080

Run it the way Pipecat recommend the source model: `temperature=0` and thinking disabled.

json
{"chat_template_kwargs": {"enable_thinking": false}}

llama.cpp resolves this model's chat format as `peg-native`; tool calls come back as proper tool_calls on /v1/chat/completions with --jinja.

⛔ You need a ROCmFPX build — stock llama.cpp will NOT load these files

ROCmFP4/ROCmFPX use ggml tensor types 100–119; upstream's table stops at 43. Build with both backends:

bash
cmake -S . -B build-hipvk -DCMAKE_BUILD_TYPE=Release \
  -DGGML_HIP=ON -DGGML_VULKAN=ON -DGGML_NATIVE=ON \
  -DVulkan_GLSLC_EXECUTABLE=/usr/bin/glslc -DAMDGPU_TARGETS=gfx1151
cmake --build build-hipvk -j
build-hipvk/bin/llama-server --list-devices   # must list ROCm0 AND Vulkan0

Head protection — why every tier uses a q8_0 head

Our usual ladder protects output.weight with `q6_K` on the 4-bit tiers. That is impossible on this model. hidden_size is 2688, and K-quants use 256-element superblocks:

2688 % 256 = 128        →  ggml.c:8024: GGML_ASSERT(start % type_traits[type].blck_size == 0) failed
[1/401] output.weight - [2688, 131072], bf16, converting to q6_K ..   SIGABRT

q8_0 uses 32-element blocks and 2688 % 32 == 0, so every tier here carries a `q8_0` headhigher precision than our usual q6_K, at a cost of roughly +88 MB on the head tensor.

We caught this as a clean natural experiment in a single run: the three q6_K-head tiers aborted in ~2 s while the q8_0-head tier built normally, same source, same binary, same moment. Note that llama-quantize --dry-run does not catch it — the dry run planned all 401 tensors and printed a clean 60247 MiB → 17223 MiB (4.58 BPW) summary. The assert only fires once real data is written.


Verification

Every tier is checked for load, coherence, and — because this is the whole point of PhoneLLM — tool calling, using the vendor-recommended mode (temperature=0, enable_thinking: false).

The tool probe is deliberately adversarial: it includes a case where the model must not call anything, and a multi-turn case, because the failure Pipecat built this model to avoid is an agent that says "yes, I've booked that" without emitting a call.

All tiers, greedy (temperature=0), enable_thinking: false, --jinja, chat format peg-native.

TierSizeLoadsCoherentTool probe
Q4_0_ROCMFP4_STRIX_LEAN15.91 GiB✅ ~10 s3/5
Q4_0_ROCMFP4_FAST15.83 GiB✅ ~10 s3/5
Q4_0_ROCMFP4_COHERENT16.91 GiB✅ ~10 s2/5
Q6_0_ROCMFPX_AGENT27.26 GiB✅ ~20 s3/5
Q8_0_ROCMFPX_AGENT30.84 GiB✅ ~25 s2/5
Q8_0_ROCMFPX30.37 GiB✅ ~20 s3/5
BF16 source (control)58.8 GiB1/5

6/6 tiers load and stay coherent. There is no precision-dependent degradation: the 30 GiB Q8 tiers score the same as the 16 GiB 4-bit tiers, and every quantised tier scores at or above the BF16 control. If quantisation were damaging tool calling, the Q8 tiers would lead. They do not — so pick on size.

Flagship detail (STRIX_LEAN), 5 adversarial cases:

PASS  booking       book_table {"name":"Chen","party_size":2,"time":"19:00"}   ← normalised "7pm" → 19:00
PASS  escalate      transfer_to_human {"reason":"Customer has called multiple times ..."}
PASS  no-tool       (correctly emitted NO call)
FAIL  availability  (no call — the one unambiguous miss)
FAIL  multiturn     check_availability {"date":"Saturday","party_size":4}

Read `3/5` carefully — the rubric is strict and opinionated. The multiturn "failure" is the model checking availability before booking, which is defensible agent behaviour; we counted it wrong because our expected answer demanded a booking. The no-tool pass matters most: the model declined to invent a call when none was warranted, which is the failure mode PhoneLLM exists to avoid. Treat these as a smoke test that the tool path survives quantisation, not as a benchmark score — for a real score use Pipecat's PhoneBench.

⚠️ An honest limitation: we could not establish a BF16 baseline on this hardware

We ran the BF16 GGUF as a control arm and it misbehaves on gfx1151 when tools are attached — the same prompt that a quantised tier answers with a correct book_table call returns, from BF16, either a degenerate repetition loop or an unrelated non-sequitur. Without tools, BF16 is coherent.

So we can report what the quantised tiers do, but we cannot publish a "delta vs BF16" the way Pipecat report NVFP4 (PhoneBench 72.06 → 71.51). Anyone quoting a quality delta for these files against BF16 on this hardware would be quoting a broken control. We have not root-caused it (candidates: gfx1151 bf16 compute, or the peg-native tool-template path); it is flagged here rather than papered over.


Reproduction block

A number without its binary is a rumour.

HostRyzen AI Max+ 395 (Strix Halo), Radeon 8060S, gfx1151, 128 GB unified
BuildROCmFPX fork @ `e7712358806055c70a9753b070202b0cc7c637e3`
GGML_HIP=ON GGML_VULKAN=ON GGML_NATIVE=ON, Release, AMDGPU_TARGETS=gfx1151
llama-server sha256e860b763d5e8f496ddb4b01236d8a8f37ae09dda456b5d076ab3abe2d9deb9fc
Sourcepipecat-ai/phonellm-alpha-1, 13 safetensors shards, 58.8 GiB
Convertedconvert_hf_to_gguf.py --outtype bf16 → 401 tensors, 63.18 GB, arch nemotron_h_moe
Quantisellama-quantize --output-tensor-type q8_0 <bf16> <out> <ftype> 12
Serve (verification)-dev ROCm0 -fa on -ngl 999 -fit off -np 1 -c 32768 -b 4096 -t 8 --jinja

Not measured

  • Perplexity (the source is a voice-agent fine-tune; wikitext PPL is a poor proxy and we would rather publish nothing than a misleading number).
  • PhoneBench — that is Pipecat's harness; we did not run it.
  • Vision — text-only model, no projector.
  • Context beyond 32768 (the source supports 262144).
  • Decode throughput per tier.

License and attribution

Released under BSD 2-Clause, matching the source. The source is itself a derivative of an NVIDIA Nemotron Open Model License work — see LICENSE_NVIDIA.txt in the upstream repo.

Acknowledgements

Daily / Pipecat for PhoneLLM and for publishing an honest PhoneBench methodology. NVIDIA for Nemotron 3 Nano and the hybrid Mamba-Transformer architecture. The ROCmFPX project for the FP4/FPX tensor types and the Strix Halo kernels that make these files possible.