petr567/LFM2.5-2.6B-Ubuntu-Strix-Halo-Vulkan-GGUF
LFM2.5-2.6B Q4KM Fast — Ubuntu Strix Halo Vulkan
This repository contains a directly runnable Q4_K_M GGUF of LiquidAI/LFM2.5-2.6B-GGUF and a validated Fast single-request Ubuntu/Vulkan profile for AMD Strix Halo systems using llama.cpp.
The model weights are not modified. The same verified GGUF is published in both paired repositories; only the tested runtime profile differs.
Fast profile: inference is accelerated relative to the same Q4_K_M baseline without the Fast runtime settings. The frozen validation gate detected no quality regression and no new failures. This is a measured result for the documented hardware, workloads, and single-request setup—not a universal guarantee for every prompt or runtime.Measured Fast result
Primary metric: wall-clock decoded tokens per second for one request, without batching. The profile validation used three workloads with five repetitions each (15 runs total, 256 generated tokens per run).
The larger Fast standard deviation reflects the deliberately mixed workload set: repetitive code benefits more than free-form editing and summarization. These are profile-validation measurements, not the pending frozen cross-machine benchmark.
Choose the matching profile
Included weight
Source revision: `b22e29ebf6249a8c9fcdda36914743e9980595c4`.
Tested setup
- Ubuntu on AMD Ryzen AI Max+ / Strix Halo
- Vulkan backend with full model offload
- 128 GiB unified memory system
- context 8,192, one parallel slot, continuous batching disabled
llama.cppVulkan server compatible with build 9994 or newer
Build llama.cpp with Vulkan
sudo apt update
sudo apt install -y git cmake build-essential libvulkan-dev glslc
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j --target llama-serverFor strict reproducibility, record the llama.cpp commit after cloning and reuse that commit for future comparisons.
Download and verify
python3 -m pip install -U huggingface_hub
MODEL_DIR="$HOME/models/LFM2.5-2.6B"
mkdir -p "$MODEL_DIR"
hf download petr567/LFM2.5-2.6B-Ubuntu-Strix-Halo-Vulkan-GGUF \
LFM2.5-2.6B-Q4_K_M.gguf \
--local-dir "$MODEL_DIR"
echo "79fdf00351b46cf26f020aead28d01889886be87c55fa0eb907e6f9b00bfee14 $MODEL_DIR/LFM2.5-2.6B-Q4_K_M.gguf" \
| sha256sum -c -Run the validated Fast Ubuntu/Vulkan profile
From the llama.cpp checkout:
MODEL_DIR="$HOME/models/LFM2.5-2.6B"
./build/bin/llama-server \
-m "$MODEL_DIR/LFM2.5-2.6B-Q4_K_M.gguf" \
--alias lfm2.5-2.6b-q4_k_m \
--host 127.0.0.1 --port 8080 \
-c 8192 -np 1 -ngl 99 \
-t 16 -tb 16 -b 2048 -ub 512 \
-fa auto \
--no-cont-batching --no-cache-prompt --cache-ram 0 \
--slot-prompt-similarity 0 --jinja --no-webui \
--spec-type ngram-simple \
--spec-ngram-simple-size-n 8 \
--spec-ngram-simple-size-m 32 \
--spec-ngram-simple-min-hits 1 \
--spec-draft-n-max 48The OpenAI-compatible endpoint is available at http://127.0.0.1:8080/v1.
curl http://127.0.0.1:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "lfm2.5-2.6b-q4_k_m",
"messages": [{"role": "user", "content": "Write a short hello-world function in Python."}],
"max_tokens": 128,
"temperature": 0.2
}'Release scope
This release contains the runnable weight and the final launch recipe. The frozen cross-machine benchmark package and its results will be attached in a later revision after verification.
Attribution and license
- Base model and GGUF: Liquid AI
- Upstream repository: LiquidAI/LFM2.5-2.6B-GGUF
- License: LFM Open License v1.0; a copy is included as `LICENSE`
The license includes a commercial-use revenue threshold. Review the included license before use or redistribution.
