CoolFace
Modelpublic

noctrex/LFM2.5-8B-A1B-MXFP4_MOE-GGUF

sourceHugging Faceupdated 4mo agoView on Hugging Face
2likes213downloads
Model Card

These are MXFP4 quantizations of the model LiquidAI / LFM2.5-8B-A1B

Quick Start

  1. 1.Download the latest release of **llama.cpp**.
  2. 2.Download your preferred model variant from below.

Which version should I choose?

All FP4 variants use MXFP4 for the MoE (Mixture of Experts) weights to keep the model efficient. I've included also a new type Q8XLMOE, that uses Q8_0 for MoE tensors and BF16 for everything else. The difference lies in how the remaining tensors are handled:

VariantQualityPerformanceMoE TensorsOther TensorsSizeRecommendation
Q8_XL_MOE⭐⭐⭐⭐⭐Variable\*Q8_0FP169.02GiBMaximum quality, uses Q8_0 instead of MXFP4 for the MoE weights.
MXFP4_MOE_BF16⭐⭐⭐Variable\*MXFP4FP165.18GiBBest for maximum accuracy; original unquantized weights.
MXFP4_MOE_F16⭐⭐FastMXFP4F165.18GiBGreat alternative if BF16 is slow on your hardware.
MXFP4_MOE⭐FastestMXFP4Q8_04.79GiBBalanced performance and memory usage.

Note: On some older architectures, BF16 may be slower than F16. Check that your GPU supports native BF16 acceleration, otherwise it would be better to get the F16 version.

Recommended parameters from LiquidAI:

  • —temperature 0.2
  • —top_p 80
  • —repetition_penalty 1.05

The chat template has been updated to fix the tool calling issues. If you don't want to download the model again, you can use the template from the parent model.