kingjones777/PhoneLLM-Alpha-1-ROCmFP4-GGUF
PhoneLLM Alpha 1 — ROCmFP4 for AMD Strix Halo (gfx1151)
The model here is not our work. PhoneLLM Alpha 1 is by [Daily](https://www.daily.co/) / the [Pipecat](https://www.pipecat.ai/) team — `pipecat-ai/phonellm-alpha-1` — a full-parameter fine-tune of [NVIDIA Nemotron 3 Nano 30B-A3B](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16). This repository adds only the ROCmFP4/ROCmFPX quantisation ladder for AMD Strix Halo and the measurements below. Go star their repo. Pipecat also ship an official NVFP4 build for NVIDIA Blackwell: `pipecat-ai/phonellm-alpha-1-nvfp4`.
A hybrid Mamba-Transformer MoE voice-agent model — 30B total, 3.5B active — quantised to run on a Ryzen AI Max+ 395 (Radeon 8060S, gfx1151). One binary, both backends: HIP (ROCm) and Vulkan are a runtime `-dev` flag, not a rebuild.
PhoneLLM is built for one job: call the right tool at the right time, with thinking disabled, at phone-call latency. That shapes how we verified it — see Verification.
Which file should I use?
Take the 15.91 GiB flagship. Across our probe the 30 GiB Q8 tiers score no better than the 16 GiB 4-bit tiers (see Verification). On a 128 GB Strix Halo that leaves real headroom to co-host your ASR and TTS models on the same box — which is the point, since PhoneLLM is the LLM stage of a voice pipeline, not a speech model (it is text-in / text-out; you still need STT and TTS).
Quick start
llama-server -m PhoneLLM-Alpha-1-Q4_0_ROCMFP4_STRIX_LEAN.gguf \
-dev ROCm0 -fa on -ngl 999 -fit off -np 1 \
-c 32768 -b 4096 -t 8 --jinja \
--host 0.0.0.0 --port 8080Run it the way Pipecat recommend the source model: `temperature=0` and thinking disabled.
{"chat_template_kwargs": {"enable_thinking": false}}llama.cpp resolves this model's chat format as `peg-native`; tool calls come back as proper tool_calls on /v1/chat/completions with --jinja.
⛔ You need a ROCmFPX build — stock llama.cpp will NOT load these files
ROCmFP4/ROCmFPX use ggml tensor types 100–119; upstream's table stops at 43. Build with both backends:
cmake -S . -B build-hipvk -DCMAKE_BUILD_TYPE=Release \
-DGGML_HIP=ON -DGGML_VULKAN=ON -DGGML_NATIVE=ON \
-DVulkan_GLSLC_EXECUTABLE=/usr/bin/glslc -DAMDGPU_TARGETS=gfx1151
cmake --build build-hipvk -j
build-hipvk/bin/llama-server --list-devices # must list ROCm0 AND Vulkan0Head protection — why every tier uses a q8_0 head
Our usual ladder protects output.weight with `q6_K` on the 4-bit tiers. That is impossible on this model. hidden_size is 2688, and K-quants use 256-element superblocks:
2688 % 256 = 128 → ggml.c:8024: GGML_ASSERT(start % type_traits[type].blck_size == 0) failed
[1/401] output.weight - [2688, 131072], bf16, converting to q6_K .. SIGABRTq8_0 uses 32-element blocks and 2688 % 32 == 0, so every tier here carries a `q8_0` head — higher precision than our usual q6_K, at a cost of roughly +88 MB on the head tensor.
We caught this as a clean natural experiment in a single run: the three q6_K-head tiers aborted in ~2 s while the q8_0-head tier built normally, same source, same binary, same moment. Note that llama-quantize --dry-run does not catch it — the dry run planned all 401 tensors and printed a clean 60247 MiB → 17223 MiB (4.58 BPW) summary. The assert only fires once real data is written.
Verification
Every tier is checked for load, coherence, and — because this is the whole point of PhoneLLM — tool calling, using the vendor-recommended mode (temperature=0, enable_thinking: false).
The tool probe is deliberately adversarial: it includes a case where the model must not call anything, and a multi-turn case, because the failure Pipecat built this model to avoid is an agent that says "yes, I've booked that" without emitting a call.
All tiers, greedy (temperature=0), enable_thinking: false, --jinja, chat format peg-native.
6/6 tiers load and stay coherent. There is no precision-dependent degradation: the 30 GiB Q8 tiers score the same as the 16 GiB 4-bit tiers, and every quantised tier scores at or above the BF16 control. If quantisation were damaging tool calling, the Q8 tiers would lead. They do not — so pick on size.
Flagship detail (STRIX_LEAN), 5 adversarial cases:
PASS booking book_table {"name":"Chen","party_size":2,"time":"19:00"} ← normalised "7pm" → 19:00
PASS escalate transfer_to_human {"reason":"Customer has called multiple times ..."}
PASS no-tool (correctly emitted NO call)
FAIL availability (no call — the one unambiguous miss)
FAIL multiturn check_availability {"date":"Saturday","party_size":4}Read `3/5` carefully — the rubric is strict and opinionated. The multiturn "failure" is the model checking availability before booking, which is defensible agent behaviour; we counted it wrong because our expected answer demanded a booking. The no-tool pass matters most: the model declined to invent a call when none was warranted, which is the failure mode PhoneLLM exists to avoid. Treat these as a smoke test that the tool path survives quantisation, not as a benchmark score — for a real score use Pipecat's PhoneBench.
⚠️ An honest limitation: we could not establish a BF16 baseline on this hardware
We ran the BF16 GGUF as a control arm and it misbehaves on gfx1151 when tools are attached — the same prompt that a quantised tier answers with a correct book_table call returns, from BF16, either a degenerate repetition loop or an unrelated non-sequitur. Without tools, BF16 is coherent.
So we can report what the quantised tiers do, but we cannot publish a "delta vs BF16" the way Pipecat report NVFP4 (PhoneBench 72.06 → 71.51). Anyone quoting a quality delta for these files against BF16 on this hardware would be quoting a broken control. We have not root-caused it (candidates: gfx1151 bf16 compute, or the peg-native tool-template path); it is flagged here rather than papered over.
Reproduction block
A number without its binary is a rumour.
Not measured
- Perplexity (the source is a voice-agent fine-tune; wikitext PPL is a poor proxy and we would rather publish nothing than a misleading number).
- PhoneBench — that is Pipecat's harness; we did not run it.
- Vision — text-only model, no projector.
- Context beyond 32768 (the source supports 262144).
- Decode throughput per tier.
License and attribution
Released under BSD 2-Clause, matching the source. The source is itself a derivative of an NVIDIA Nemotron Open Model License work — see LICENSE_NVIDIA.txt in the upstream repo.
- Model: `pipecat-ai/phonellm-alpha-1` — Daily / Pipecat.
- Base: `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16` — NVIDIA.
- This repository contributes only the ROCmFP4/ROCmFPX quantisation ladder and the measurements above.
Acknowledgements
Daily / Pipecat for PhoneLLM and for publishing an honest PhoneBench methodology. NVIDIA for Nemotron 3 Nano and the hybrid Mamba-Transformer architecture. The ROCmFPX project for the FP4/FPX tensor types and the Strix Halo kernels that make these files possible.
