noctrex/LFM2-8B-A1B-MXFP4_MOE-GGUF
179
These are MXFP4 quantizations of the model LiquidAI / LFM2-8B-A1B
Quick Start
- Download the latest release of llama.cpp.
- Download your preferred model variant from below.
Which version should I choose?
All FP4 variants use MXFP4 for the MoE (Mixture of Experts) weights to keep the model efficient. I've included also a new type Q8XLMOE, that uses Q8 for MoE tensors and BF16 for everything else. The difference lies in how the remaining tensors are handled:
Note: On some older architectures, BF16 may be slower than F16. Check that your GPU supports native BF16
Recommended parameters from LiquidAI:
- temperature 0.3
- min_p 0.15
- repetition_penalty 1.05
