amd/MiniMax-M2.1-MXFP4
11.5k
1---2base_model:3- MiniMaxAI/MiniMax-M2.14language:5- en6library_name: transformers7license: other8license_name: modified-mit9license_link: https://github.com/MiniMax-AI/MiniMax-M2.1/blob/main/LICENSE10---11 12# Model Overview13 14- **Model Architecture:** MiniMaxM2ForCausalLM15 - **Input:** Text16 - **Output:** Text17- **Supported Hardware Microarchitecture:** AMD MI300 MI350/MI35518- **ROCm**: 7.019- **PyTorch**: 2.8.020- **Transformers**: 4.57.121- **Operating System(s):** Linux22- **Inference Engine:** [SGLang](https://docs.sglang.ai/)/[vLLM](https://docs.vllm.ai/en/latest/)23- **Model Optimizer:** [AMD-Quark](https://quark.docs.amd.com/latest/index.html) (v0.11)24 - **Weight quantization:** OCP MXFP4, Static25 - **Activation quantization:** OCP MXFP4, Dynamic26 27 28# Model Quantization29 30The model was quantized from [QuixiAI/MiniMax-M2.1-bf16](https://huggingface.co/QuixiAI/MiniMax-M2.1-bf16) using [AMD-Quark](https://quark.docs.amd.com/latest/index.html). The weights are quantized to MXFP4 and activations are quantized to MXFP4.31 32 33**Quantization scripts:**34```35cd Quark/examples/torch/language_modeling/llm_ptq/36export exclude_layers="lm_head *block_sparse_moe.gate* *self_attn*"37python3 quantize_quark.py --model_dir $MODEL_DIR \38 --quant_scheme mxfp4 \39 --num_calib_data 128 \40 --exclude_layers $exclude_layers \41 --skip_evaluation \42 --multi_gpu \43 --trust_remote_code \44 --model_export hf_format \45 --output_dir $output_dir46```47For further details or issues, please refer to the AMD-Quark documentation or contact the respective developers.48 49# Evaluation50The model was evaluated on gsm8k benchmarks using the [vllm](https://github.com/vllm-project/vllm/tree/v0.13.0) framework.51 52### Accuracy53 54<table>55 <tr>56 <td><strong>Benchmark</strong>57 </td>58 <td><strong>QuixiAI/MiniMax-M2.1-bf16 </strong>59 </td>60 <td><strong>amd/MiniMax-M2.1-MXFP4(this model)</strong>61 </td>62 <td><strong>Recovery</strong>63 </td>64 </tr>65 <tr>66 <td>gsm8k (flexible-extract) 67 </td>68 <td>0.935669 </td>70 <td>0.934871 </td>72 <td>99.91%73 </td>74 </tr>75</table>76 77 78### Reproduction79 80The GSM8K results were obtained using the vLLM framework, based on the Docker image `rocm/vllm-dev:nightly_main_20260211`, and vLLM is installed inside the container.81 82#### Preparation in container83 84To download the evaluation script, reinstallation is not required.85```86# Install vLLM code repo87git clone https://github.com/vllm-project/vllm.git88cd vllm89git checkout v0.13.090cd ..91```92 93#### Launching server94```95VLLM_ROCM_USE_AITER=1 \96VLLM_DISABLE_COMPILE_CACHE=1 \97vllm serve "$MODEL" \98 --tensor-parallel-size 4 \99 --trust-remote-code \100 --max-model-len 32768 \101 --port 8899 102```103 104 105#### Evaluating model in a new terminal106```107python vllm/tests/evals/gsm8k/gsm8k_eval.py --host http://127.0.0.1 --port 8899 --num-questions 1000 --save-results logs108```109 110 111# License112Modifications Copyright(c) 2026 Advanced Micro Devices, Inc. All rights reserved.