FadedRedStar/LFM2.5-8B-A1B-heretic-GGUF
๐ค LFM2.5-8B-A1B-heretic โ GGUF
This repository hosts GGUF weights for LFM2.5-8B-A1B-heretic, quantized from the source floating-point tensors provided by coder3101/LFM2.5-8B-A1B-heretic.
๐ Sister Repository: Check out the Imatrix Sister Repository for enhanced precision at lower bit fractions.
[!NOTE] If you plan on using 4-bit or 5-bit variants, consider the imatrix sister repository instead โ importance matrix calibration improves logic retention at those bit depths. This repository is best suited if you want the near-lossless Q8_0 build.โน๏ธ Model Profile & Core Features
LFM2.5-8B-A1B is a text-only model from Liquid AI's Liquid Foundation Model 2.5 series, designed for on-device deployment. It uses a hybrid architecture with 24 layers โ 18 double-gated LIV (Liquid, Input-adaptive, Value-selective) convolution layers plus 6 GQA (Grouped Query Attention) layers โ activating only approximately 1.5B parameters per forward pass out of 8.3B total. This delivers fastest-in-class throughput at its size on both CPU and GPU, with day-one support for llama.cpp, MLX, vLLM, and SGLang. The model is a reasoning model: it produces a chain-of-thought before its final answer, and is tuned for complex instruction following, tool calling, and chained agentic task execution.
The heretic suffix denotes post-processing via the [Heretic v1.2.0 Arbitrary-Rank Ablation (ARA)](https://github.com/p-e-w/heretic) method with row-norm preservation performed by coder3101, which removes refusal conditioning at multiple tensor ranks while maintaining the model's instruction-following and planning capabilities.
๐ Technical Specifications
๐ ๏ธ Heretic Overrides (ARA)
๐ Refusal Bypass Metrics
[!NOTE] The metrics below are self-reported by the original model author (coder3101) and have not been independently reproduced.
๐งฎ Numerical & Tensor Formats
๐ฆ Available Model Files
Main model weights | Filename | Quantization | llama.cpp Build | Size | Download | |---|---|---|---|---| | LFM2.5-8B-A1B-heretic-Q4_K_M.gguf | Q4_K_M | b9803 | 4.80 GB | ๐ฅ Download | | LFM2.5-8B-A1B-heretic-Q5_K_M.gguf | Q5_K_M | b9870 | 5.62 GB | ๐ฅ Download | | LFM2.5-8B-A1B-heretic-Q8_0.gguf | Q8_0 | b9870 | 8.39 GB | ๐ฅ Download |
๐๏ธ Component Pairing Guide
Download exactly one main weights file:
- `Q4_K_M`: Balanced 4-bit format suitable for most everyday use.
- `Q5_K_M`: Higher-fidelity mid-range format recommended as a general default.
- `Q8_0`: Near-lossless 8-bit format for when memory is not a constraint.
โก Deployment & Execution Commands
[!NOTE] Liquid AI recommends the following generation parameters for best results:temperature: 0.2,top_k: 80,repetition_penalty: 1.05.
[!NOTE] This model emits reasoning content before its final answer. If you require a clean final answer only, parse the output accordingly rather than expecting a single direct response.
[!TIP] Swap the -m filename below for either quantized file depending on your size/quality trade-off preference.llama.cpp CLI
./llama-cli \
-m LFM2.5-8B-A1B-heretic-Q4_K_M.gguf \
-c 8192 \
-ngl 99 \
--temp 0.2 \
--top-k 80 \
--repeat-penalty 1.05 \
-p "<|im_start|>system\nYou are a helpful and precise assistant capable of using tools and following complex instructions.<|im_end|>\n<|im_start|>user\nBreak down the following task and execute it step by step: summarise this document and list action items.<|im_end|>\n<|im_start|>assistant\n"OpenAI-Compatible API Server
./llama-server \
--host 0.0.0.0 \
--port 8080 \
-m LFM2.5-8B-A1B-heretic-Q4_K_M.gguf \
-c 16384 \
-ngl 99 \
--flash-attn๐ฌ Chat Templates & Prompt Design (ChatML)
<|im_start|>system
You are a capable assistant. Follow instructions precisely.<|im_end|>
<|im_start|>user
Your task or query here.<|im_end|>
<|im_start|>assistantโ ๏ธ Safety & Operational Notes
- This model is abliterated and will generate content that standard aligned models refuse. Use responsibly and in compliance with applicable laws.
- This is a text-only model โ it has no vision encoder and cannot process images.
- The LIV architecture activates only ~1.5B parameters per token, making it significantly faster to run than the total parameter count implies.
- For long-context workloads, set
-cup to 131072 as needed. - Liquid AI shipped a tokenizer fix for tool-calling after this model's initial release; if you encounter malformed tool-call output, verify your llama.cpp build includes this fix.
- For better output quality at this quantization level, consider the imatrix variant in the companion repository.
