CoolFace
Modelpublic

sparkarena/Minimax-M3-v0-NVFP4-REAP50

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
14likes129downloads
Model Card

NOTICE

This is an experimental quantization of MiniMax M3 to NVFP4 for use on DGX Spark.

On top of that, REAP 50% to make it fix on 2x DGX Sparks instead of 4x.

NVFP4 uses w1/3 scales, so need either: https://github.com/scitrera/sglang/tree/nvfp4-w13-scale-normalization (or possibly sglang PR#27588) in place to properly use the calibration from this model.

The calibration is still a work in progress (hence "v0" model). Updates are planned to improve performance.

Re: REAP50, I've tested that it's coherent, but I haven't exhaustively tested the quality degradation associated with REAP50. YMMV.

Run with sparkrun; part of Spark Arena

https://sparkrun.dev https://spark-arena.com

To run with sparkrun on 2x DGX Spark Nodes:

sparkrun run @experimental/minimax-m3-v0-nvfp4-2x-reap50

<div align="center"> <img width="60%" src="figures/logo.svg" alt="MiniMax"> </div> <hr>

<p align="center"> <a href="https://agent.minimax.io/" target="blank"><img src="https://img.shields.io/badge/MiniMax%20Agent-FF6C37?style=for-the-badge&logo=minimax&logoColor=white" alt="MiniMax Agent"></a> <a href="https://platform.minimax.io/docs/guides/text-generation" target="blank"><img src="https://img.shields.io/badge/API-FF6C37?style=for-the-badge&logo=minimax&logoColor=white" alt="API"></a> <a href="https://www.minimax.io" target="blank"><img src="https://img.shields.io/badge/MiniMax%20Website-FF6C37?style=for-the-badge&logo=minimax&logoColor=white" alt="MiniMax Website"></a> <br> <a href="https://platform.minimaxi.com/docs/faq/contact-us" target="blank"><img src="https://img.shields.io/badge/WeChat-07C160?style=for-the-badge&logo=wechat&logoColor=white" alt="WeChat"></a> <a href="https://discord.com/invite/DPC4AHFCBw" target="blank"><img src="https://img.shields.io/badge/Discord-5865F2?style=for-the-badge&logo=discord&logoColor=white" alt="Discord"></a> <a href="https://huggingface.co/MiniMaxAI" target="blank"><img src="https://img.shields.io/badge/Hugging%20Face-FFD21E?style=for-the-badge&logo=huggingface&logoColor=black" alt="Hugging Face"></a> <a href="https://github.com/MiniMax-AI/MiniMax-M3" target="blank"><img src="https://img.shields.io/badge/GitHub-181717?style=for-the-badge&logo=github&logoColor=white" alt="GitHub"></a> <a href="https://arxiv.org/abs/2606.13392" target="blank"><img src="https://img.shields.io/badge/arXiv-2606.13392-B31B1B?style=for-the-badge&logo=arxiv&logoColor=white" alt="arXiv Paper"></a> <a href="https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/LICENSE" target="_blank"><img src="https://img.shields.io/badge/LICENSE-4CAF50?style=for-the-badge&logo=creativecommons&logoColor=white" alt="LICENSE"></a> </p>

MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.

Highlights:

  • —Native Multimodality: M3 undergoes mixed-modality training from the very first step, enabling deeper semantic fusion across text, image, and video.
  • —Context Scaling via Sparse Attention: M3 introduces MiniMax Sparse Attention (MSA) to improve long context efficiency. M3 delivers 9× prefill and 15× decode speedups compared to M2 at 1M context, reducing per-token compute to 1/20.
  • —Coding & Cowork Capability: M3 achieves frontier-level performance across long-horizon agentic benchmarks, excelling in both coding and cowork.

<p align="center"> <img width="100%" src="figures/benchmark.jpeg"> </p>

MiniMax Sparse Attention (MSA)

M3 is powered by **MiniMax Sparse Attention (MSA)**, a high-performance sparse attention operator designed for million-token contexts. Compared with GQA, MSA dramatically reduces the attention compute and memory footprint while preserving model quality.

<p align="center"> <img width="100%" src="figures/efficiencygqavs_msa.png" alt="GQA vs MSA Efficiency Comparison"> </p>

📄 Read the technical report: arXiv:2606.13392 · Hugging Face Papers

How to Use

M3 supports two reasoning modes:

  • —thinking — for complex reasoning, agentic tasks, and long-horizon collaboration.
  • —non-thinking — for latency-sensitive scenarios such as chat and code completion.

Local Deployment

Download the model:

bash
hf download MiniMaxAI/MiniMax-M3 --local-dir MiniMax-M3

We recommend the following inference frameworks (listed alphabetically) to serve the model:

Inference Parameters

We recommend the following parameters for best performance: temperature=1.0, top_p=0.95, top_k=40.

Contact Us

Contact us at model@minimax.io.