CoolFace
Modelpublic

petr567/LFM2.5-2.6B-Ubuntu-Strix-Halo-Vulkan-GGUF

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
1likes59downloads
Model Card

LFM2.5-2.6B Q4KM Fast — Ubuntu Strix Halo Vulkan

This repository contains a directly runnable Q4_K_M GGUF of LiquidAI/LFM2.5-2.6B-GGUF and a validated Fast single-request Ubuntu/Vulkan profile for AMD Strix Halo systems using llama.cpp.

The model weights are not modified. The same verified GGUF is published in both paired repositories; only the tested runtime profile differs.

Fast profile: inference is accelerated relative to the same Q4_K_M baseline without the Fast runtime settings. The frozen validation gate detected no quality regression and no new failures. This is a measured result for the documented hardware, workloads, and single-request setup—not a universal guarantee for every prompt or runtime.

Measured Fast result

Primary metric: wall-clock decoded tokens per second for one request, without batching. The profile validation used three workloads with five repetitions each (15 runs total, 256 generated tokens per run).

WorkloadBaseline, tok/sFast, tok/sFast vs baseline
Code copy108.97231.562.125× (+112.5%)
Editorial rewrite106.75127.551.195× (+19.5%)
Technical summary105.86187.061.767× (+76.7%)
All 15 runs, mean ± SD107.20 ± 1.53182.06 ± 44.301.698× (+69.8%)
Independent quality gateBaselineFastRegression
Passed tasks9/129/12None measured

The larger Fast standard deviation reflects the deliberately mixed workload set: repetitive code benefits more than free-form editing and summarization. These are profile-validation measurements, not the pending frozen cross-machine benchmark.

Choose the matching profile

PlatformHardware/backendRepository
Windows 11NVIDIA RTX / CUDALFM2.5-2.6B-Windows-RTX-CUDA-GGUF
UbuntuAMD Strix Halo / VulkanThis repository

Included weight

FileQuantizationSizeSHA-256
`LFM2.5-2.6B-Q4_K_M.gguf`Q4KM1,674,454,848 bytes (1.56 GiB)79fdf00351b46cf26f020aead28d01889886be87c55fa0eb907e6f9b00bfee14

Source revision: `b22e29ebf6249a8c9fcdda36914743e9980595c4`.

Tested setup

  • —Ubuntu on AMD Ryzen AI Max+ / Strix Halo
  • —Vulkan backend with full model offload
  • —128 GiB unified memory system
  • —context 8,192, one parallel slot, continuous batching disabled
  • —llama.cpp Vulkan server compatible with build 9994 or newer

Build llama.cpp with Vulkan

bash
sudo apt update
sudo apt install -y git cmake build-essential libvulkan-dev glslc

git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j --target llama-server

For strict reproducibility, record the llama.cpp commit after cloning and reuse that commit for future comparisons.

Download and verify

bash
python3 -m pip install -U huggingface_hub

MODEL_DIR="$HOME/models/LFM2.5-2.6B"
mkdir -p "$MODEL_DIR"

hf download petr567/LFM2.5-2.6B-Ubuntu-Strix-Halo-Vulkan-GGUF \
  LFM2.5-2.6B-Q4_K_M.gguf \
  --local-dir "$MODEL_DIR"

echo "79fdf00351b46cf26f020aead28d01889886be87c55fa0eb907e6f9b00bfee14  $MODEL_DIR/LFM2.5-2.6B-Q4_K_M.gguf" \
  | sha256sum -c -

Run the validated Fast Ubuntu/Vulkan profile

From the llama.cpp checkout:

bash
MODEL_DIR="$HOME/models/LFM2.5-2.6B"

./build/bin/llama-server \
  -m "$MODEL_DIR/LFM2.5-2.6B-Q4_K_M.gguf" \
  --alias lfm2.5-2.6b-q4_k_m \
  --host 127.0.0.1 --port 8080 \
  -c 8192 -np 1 -ngl 99 \
  -t 16 -tb 16 -b 2048 -ub 512 \
  -fa auto \
  --no-cont-batching --no-cache-prompt --cache-ram 0 \
  --slot-prompt-similarity 0 --jinja --no-webui \
  --spec-type ngram-simple \
  --spec-ngram-simple-size-n 8 \
  --spec-ngram-simple-size-m 32 \
  --spec-ngram-simple-min-hits 1 \
  --spec-draft-n-max 48

The OpenAI-compatible endpoint is available at http://127.0.0.1:8080/v1.

bash
curl http://127.0.0.1:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "lfm2.5-2.6b-q4_k_m",
    "messages": [{"role": "user", "content": "Write a short hello-world function in Python."}],
    "max_tokens": 128,
    "temperature": 0.2
  }'

Release scope

This release contains the runnable weight and the final launch recipe. The frozen cross-machine benchmark package and its results will be attached in a later revision after verification.

Attribution and license

The license includes a commercial-use revenue threshold. Review the included license before use or redistribution.