FadedRedStar/LFM2.5-8B-A1B-heretic-imatrix-GGUF
๐ค LFM2.5-8B-A1B-heretic โ Importance Matrix GGUF
This repository hosts importance-matrix (imatrix) optimized GGUF weights, available in multiple quantization formats, for LFM2.5-8B-A1B-heretic, quantized from the source floating-point tensors provided by coder3101/LFM2.5-8B-A1B-heretic.
๐ Sister Repository: Check out the Standard GGUF Sister Repository for uncalibrated and full 8-bit precision options.
๐ฏ Matrix-Weighted Calibration (Imatrix)
An Importance Matrix (imatrix) calculation tracks activations across network layers using a calibration sequence, then weights the quantization process to preserve the parameters that matter most for output quality โ improving fidelity at low bit depths.
โก๏ธ Calibration dataset: Bartowski's calibration_datav5.txt.
[!NOTE] `IQ4_NL` is included because the matrix enables a non-linear 4-bit format that outperforms standard linear 4-bit quantization. `Q8_0` is absent because 8-bit quantization already introduces near-zero degradation, making calibration unnecessary โ see the standard sister repository for that variant.
โน๏ธ Model Profile & Core Features
LFM2.5-8B-A1B is a text-only model from Liquid AI's Liquid Foundation Model 2.5 series, designed for on-device deployment. It uses a hybrid architecture with 24 layers โ 18 double-gated LIV (Liquid, Input-adaptive, Value-selective) convolution layers plus 6 GQA (Grouped Query Attention) layers โ activating only approximately 1.5B parameters per forward pass out of 8.3B total. This delivers fastest-in-class throughput at its size on both CPU and GPU, with day-one support for llama.cpp, MLX, vLLM, and SGLang. The model is a reasoning model: it produces a chain-of-thought before its final answer, and is tuned for complex instruction following, tool calling, and chained agentic task execution.
The heretic suffix denotes post-processing via the [Heretic v1.2.0 Arbitrary-Rank Ablation (ARA)](https://github.com/p-e-w/heretic) method with row-norm preservation performed by coder3101, which removes refusal conditioning at multiple tensor ranks while maintaining the model's instruction-following and planning capabilities.
๐ Technical Specifications
๐ ๏ธ Heretic Overrides (ARA)
๐ Refusal Bypass Metrics
[!NOTE] The metrics below are self-reported by the original model author (coder3101) and have not been independently reproduced.
๐งฎ Numerical & Tensor Formats
๐ฆ Available Model Files
Main model weights | Filename | Quantization | llama.cpp Build | Size | Download | |---|---|---|---|---| | LFM2.5-8B-A1B-heretic-IQ4_NL-imatrix.gguf | IQ4_NL | b9843 | 4.51 GB | ๐ฅ Download | | LFM2.5-8B-A1B-heretic-Q4_K_M-imatrix.gguf | Q4_K_M | b9803 | 4.80 GB | ๐ฅ Download | | LFM2.5-8B-A1B-heretic-Q5_K_M-imatrix.gguf | Q5_K_M | b9870 | 5.62 GB | ๐ฅ Download |
๐๏ธ Component Pairing Guide
Download exactly one main weights file:
- `IQ4_NL`: Non-linear 4-bit format, best choice for constrained memory when imatrix calibration is present.
- `Q4_K_M`: Balanced 4-bit format suitable for most everyday use.
- `Q5_K_M`: Higher-fidelity mid-range format recommended as a general default.
โก Deployment & Execution Commands
[!NOTE] Liquid AI recommends the following generation parameters for best results:temperature: 0.2,top_k: 80,repetition_penalty: 1.05.
[!NOTE] This model emits reasoning content before its final answer. If you require a clean final answer only, parse the output accordingly rather than expecting a single direct response.
[!TIP] Swap the -m filename below for either quantized file depending on your size/quality trade-off preference.llama.cpp CLI
./llama-cli \
-m LFM2.5-8B-A1B-heretic-IQ4_NL-imatrix.gguf \
-c 8192 \
-ngl 99 \
--temp 0.2 \
--top-k 80 \
--repeat-penalty 1.05 \
-p "<|im_start|>system\nYou are a helpful and precise assistant capable of using tools and following complex instructions.<|im_end|>\n<|im_start|>user\nBreak down the following task and execute it step by step: summarise this document and list action items.<|im_end|>\n<|im_start|>assistant\n"OpenAI-Compatible API Server
./llama-server \
--host 0.0.0.0 \
--port 8080 \
-m LFM2.5-8B-A1B-heretic-IQ4_NL-imatrix.gguf \
-c 16384 \
-ngl 99 \
--flash-attn๐ฌ Chat Templates & Prompt Design (ChatML)
<|im_start|>system
You are a capable assistant. Follow instructions precisely.<|im_end|>
<|im_start|>user
Your task or query here.<|im_end|>
<|im_start|>assistantโ ๏ธ Safety & Operational Notes
- This model is abliterated and will generate content that standard aligned models refuse. Use responsibly and in compliance with applicable laws.
- This is a text-only model โ it has no vision encoder and cannot process images.
- The LIV architecture activates only ~1.5B parameters per token, making it significantly faster to run than the total parameter count implies.
- For long-context workloads, set
-cup to 131072 as needed. - Liquid AI shipped a tokenizer fix for tool-calling after this model's initial release; if you encounter malformed tool-call output, verify your llama.cpp build includes this fix.
- Imatrix calibration improves perplexity recovery compared to non-imatrix quantization, particularly on low-frequency tokens.
- IQ4NL produces a smaller file than Q4K_M and tends to run faster on CPU and ARM devices; imatrix calibration narrows the quality gap between the two formats considerably.
