CoolFace
Modelpublic

sparkarena/Minimax-M3-v0-NVFP4

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
15likes152downloads
Model Card

The First Published NVFP4 Quantization of MiniMax M3

NOTICE

This is an experimental quantization of MiniMax M3 to NVFP4 for use on DGX Spark. (Note: This quantization is not DGX Spark only.)

For DGX Spark Users: Run with sparkrun; part of Spark Arena

https://sparkrun.dev https://spark-arena.com

To run with sparkrun on 4x DGX Spark Nodes:

sparkrun run @experimental/minimax-m3-v0-nvfp4-4x

For RTX Pro 6000 Users (or DGX Spark Users who don't want to use sparkrun): You can run this using the custom sglang container:

docker pull scitrera/dgx-spark-sglang-mm:v0

(Container build is multi-arch so it can be used for x86 and ARM)

Reference settings can be derived from the sparkrun recipe: https://github.com/spark-arena/recipe-registry/blob/main/experimental-recipes/minimax-m3/minimax-m3-v0-nvfp4-4x.yaml

Happy Coding! Let's go!


<div align="center"> <img width="60%" src="figures/logo.svg" alt="MiniMax"> </div> <hr>

<p align="center"> <a href="https://agent.minimax.io/" target="blank"><img src="https://img.shields.io/badge/MiniMax%20Agent-FF6C37?style=for-the-badge&logo=minimax&logoColor=white" alt="MiniMax Agent"></a> <a href="https://platform.minimax.io/docs/guides/text-generation" target="blank"><img src="https://img.shields.io/badge/API-FF6C37?style=for-the-badge&logo=minimax&logoColor=white" alt="API"></a> <a href="https://www.minimax.io" target="blank"><img src="https://img.shields.io/badge/MiniMax%20Website-FF6C37?style=for-the-badge&logo=minimax&logoColor=white" alt="MiniMax Website"></a> <br> <a href="https://platform.minimaxi.com/docs/faq/contact-us" target="blank"><img src="https://img.shields.io/badge/WeChat-07C160?style=for-the-badge&logo=wechat&logoColor=white" alt="WeChat"></a> <a href="https://discord.com/invite/DPC4AHFCBw" target="blank"><img src="https://img.shields.io/badge/Discord-5865F2?style=for-the-badge&logo=discord&logoColor=white" alt="Discord"></a> <a href="https://huggingface.co/MiniMaxAI" target="blank"><img src="https://img.shields.io/badge/Hugging%20Face-FFD21E?style=for-the-badge&logo=huggingface&logoColor=black" alt="Hugging Face"></a> <a href="https://github.com/MiniMax-AI/MiniMax-M3" target="blank"><img src="https://img.shields.io/badge/GitHub-181717?style=for-the-badge&logo=github&logoColor=white" alt="GitHub"></a> <a href="https://arxiv.org/abs/2606.13392" target="blank"><img src="https://img.shields.io/badge/arXiv-2606.13392-B31B1B?style=for-the-badge&logo=arxiv&logoColor=white" alt="arXiv Paper"></a> <a href="https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/LICENSE" target="_blank"><img src="https://img.shields.io/badge/LICENSE-4CAF50?style=for-the-badge&logo=creativecommons&logoColor=white" alt="LICENSE"></a> </p>

MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.

Highlights:

  • —Native Multimodality: M3 undergoes mixed-modality training from the very first step, enabling deeper semantic fusion across text, image, and video.
  • —Context Scaling via Sparse Attention: M3 introduces MiniMax Sparse Attention (MSA) to improve long context efficiency. M3 delivers 9× prefill and 15× decode speedups compared to M2 at 1M context, reducing per-token compute to 1/20.
  • —Coding & Cowork Capability: M3 achieves frontier-level performance across long-horizon agentic benchmarks, excelling in both coding and cowork.

<p align="center"> <img width="100%" src="figures/benchmark.jpeg"> </p>

MiniMax Sparse Attention (MSA)

M3 is powered by **MiniMax Sparse Attention (MSA)**, a high-performance sparse attention operator designed for million-token contexts. Compared with GQA, MSA dramatically reduces the attention compute and memory footprint while preserving model quality.

<p align="center"> <img width="100%" src="figures/efficiencygqavs_msa.png" alt="GQA vs MSA Efficiency Comparison"> </p>

📄 Read the technical report: arXiv:2606.13392 · Hugging Face Papers

How to Use

M3 supports two reasoning modes:

  • —thinking — for complex reasoning, agentic tasks, and long-horizon collaboration.
  • —non-thinking — for latency-sensitive scenarios such as chat and code completion.

Local Deployment

Download the model:

bash
hf download MiniMaxAI/MiniMax-M3 --local-dir MiniMax-M3

We recommend the following inference frameworks (listed alphabetically) to serve the model:

Inference Parameters

We recommend the following parameters for best performance: temperature=1.0, top_p=0.95, top_k=40.

Contact Us

Contact us at model@minimax.io.