JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-MXFP4_MOE
Mellum2 Instruct — GGUF (MXFP4_MOE)
This repository contains a GGUF MXFP4_MOE quantization of `JetBrains/Mellum2-12B-A2.5B-Instruct`, ready to run with `llama.cpp`, Ollama, LM Studio, and other GGUF-compatible runtimes.
This quantization (MXFP4_MOE): MXFP4 microscaling 4-bit applied to the MoE expert tensors. Smallest footprint, with a quality cost (KLD ~0.166, 84% top-token agreement).
Mellum 2 Instruct is a Mixture-of-Experts assistant model (64 experts, 8 activated per token, 131,072-token context) that answers directly, without an externalized chain of thought. For the full model description, evaluation results, and architecture details, see the original model card: [JetBrains/Mellum2-12B-A2.5B-Instruct](https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Instruct).
Available quantizations
KL divergence and top-token agreement are measured against the BF16 logits on Wikitext-2 (n_ctx=512); lower KLD / higher agreement means closer to the unquantized model. (Perplexity is omitted here — it is unreliable for instruction-tuned models on Wikitext-2, which is out of distribution.)
Download
hf download JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-MXFP4_MOE Mellum2-12B-A2.5B-Instruct-MXFP4_MOE.gguf --local-dir .Run with llama.cpp
# Pull and serve in one step (downloads the GGUF automatically)
llama-server -hf JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-MXFP4_MOE \
--ctx-size 131072 \
--temp 0.6 --top-p 0.95 --top-k 20
# Or run a one-off prompt with a local file
llama-cli -m Mellum2-12B-A2.5B-Instruct-MXFP4_MOE.gguf \
--ctx-size 131072 \
--temp 0.6 --top-p 0.95 --top-k 20 \
-p "Write a Python function to reverse a string."The server exposes an OpenAI-compatible API on http://localhost:8080/v1:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8080/v1", api_key="llama.cpp")
chat_response = client.chat.completions.create(
model="JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-MXFP4_MOE",
messages=[
{"role": "user", "content": "Write a Python function to reverse a string."},
],
max_tokens=81920,
temperature=0.6,
top_p=0.95,
extra_body={"top_k": 20},
)
print(chat_response.choices[0].message.content)Run with Ollama
ollama run hf.co/JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-MXFP4_MOELicense
Released under the Apache 2.0 license.
For the full model card, evaluation results, and architecture details, refer to the original model: [JetBrains/Mellum2-12B-A2.5B-Instruct](https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Instruct).
