mlx-community/K2-Horizon-MoVA-36B-A4B-oQ4e
K2-Horizon-MoVA-36B-A4B oQ4e
oQ4e (imatrix-enhanced mixed-precision 4-bit) conversion of IFM/K2-Horizon-MoVA-36B-A4B, a sparse Mixture-of-Experts model with Mixture-of-Values attention (36B total / 4B active parameters, 512K context).
Upstream model: IFM/K2-Horizon-MoVA-36B-A4B by the IFM Team, released under Apache-2.0.
Conversion: Quantized to MLX format using Hermes Agent with mlx-lm and oMLX.
Quickstart
pip install -U mlx-lm
python3 -m mlx_lm.generate \
--model hermitdave/K2-Horizon-MoVA-36B-A4B-oQ4e \
--prompt "Explain why long-context evaluation is difficult." \
--max-tokens 512 --temp 1.0 --top-p 0.95Reasoning
K2-Horizon is a reasoning model. Always use reasoning_effort="high" for best results:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="hermitdave/K2-Horizon-MoVA-36B-A4B-oQ4e",
messages=[{"role": "user", "content": "Explain quantum entanglement."}],
extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print("Reasoning:", getattr(response.choices[0].message, "reasoning_content", None))
print("Answer:", response.choices[0].message.content)Benchmark Results
Leaderboard Benchmarks
Scores in %. See model card for full results.
Head-to-Head Comparison vs Qwen 3.6 35B A3B
Independent benchmark comparison — raw scores collected via BenchLocal, report compiled by Hermes Agent. See full report.
K2 wins 7/8 benchmarks. Dominates extraction, formatting, tool calling, and safety. Qwen wins only Reasoning & Maths by 2 points.
Throughput Comparison (oMLX)
Qwen is ~1.8x faster due to 3B active parameters (vs 4B) + MTP (multi-token prediction). K2 trades throughput for quality.
Verdict by Dimension
Bottom line: For local agent pipelines where tool reliability and structured output matter: K2 Horizon. For high-volume batch inference where throughput dominates: Qwen 3.6.
oMLX Patch
K2-Horizon requires oMLX v0.6.4+ with the K2-Horizon support patch (PR #3441). This patch adds:
k2_horizonmodel type support- Reasoning content handling (
<ifm|think>tags) - Tool call parsing (plain text and XML formats)
- Multi-turn conversation support
Without this patch, oMLX will refuse to load K2-Horizon models with ValueError: Model type k2_horizon not supported.
Chat Template
K2-Horizon uses IFM's custom chat template with reasoning and tool calling support. Key tags:
All tags are automatically stripped by oMLX before responses reach users.
Citation
@misc{k2horizon2026,
title = {Introducing K2 Horizon: Frontier Performance, Radically Open},
author = {{IFM Team}},
year = {2026},
url = {https://ifm.ai/blog/k2/},
}License
Apache-2.0 (same as upstream).
