CoolFace
Modelpublic

amd/MiniMax-M2.1-MXFP4

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
1likes1.5kdownloads
README.md112 linesDownload Raw Back to root
1---2base_model:3- MiniMaxAI/MiniMax-M2.14language:5- en6library_name: transformers7license: other8license_name: modified-mit9license_link: https://github.com/MiniMax-AI/MiniMax-M2.1/blob/main/LICENSE10---11 12# Model Overview13 14- **Model Architecture:** MiniMaxM2ForCausalLM15  - **Input:** Text16  - **Output:** Text17- **Supported Hardware Microarchitecture:** AMD MI300 MI350/MI35518- **ROCm**: 7.019- **PyTorch**: 2.8.020- **Transformers**: 4.57.121- **Operating System(s):** Linux22- **Inference Engine:** [SGLang](https://docs.sglang.ai/)/[vLLM](https://docs.vllm.ai/en/latest/)23- **Model Optimizer:** [AMD-Quark](https://quark.docs.amd.com/latest/index.html) (v0.11)24  - **Weight quantization:** OCP MXFP4, Static25  - **Activation quantization:** OCP MXFP4, Dynamic26 27 28# Model Quantization29 30The model was quantized from [QuixiAI/MiniMax-M2.1-bf16](https://huggingface.co/QuixiAI/MiniMax-M2.1-bf16) using [AMD-Quark](https://quark.docs.amd.com/latest/index.html). The weights are quantized to MXFP4 and activations are quantized to MXFP4.31 32 33**Quantization scripts:**34```35cd Quark/examples/torch/language_modeling/llm_ptq/36export exclude_layers="lm_head *block_sparse_moe.gate* *self_attn*"37python3 quantize_quark.py --model_dir $MODEL_DIR \38                          --quant_scheme mxfp4 \39                          --num_calib_data 128 \40                          --exclude_layers $exclude_layers \41                          --skip_evaluation \42                          --multi_gpu  \43                          --trust_remote_code \44                          --model_export hf_format \45                          --output_dir $output_dir46```47For further details or issues, please refer to the AMD-Quark documentation or contact the respective developers.48 49# Evaluation50The model was evaluated on gsm8k benchmarks using the [vllm](https://github.com/vllm-project/vllm/tree/v0.13.0) framework.51 52### Accuracy53 54<table>55  <tr>56   <td><strong>Benchmark</strong>57   </td>58   <td><strong>QuixiAI/MiniMax-M2.1-bf16 </strong>59   </td>60   <td><strong>amd/MiniMax-M2.1-MXFP4(this model)</strong>61   </td>62   <td><strong>Recovery</strong>63   </td>64  </tr>65  <tr>66   <td>gsm8k (flexible-extract) 67   </td>68   <td>0.935669   </td>70   <td>0.934871   </td>72   <td>99.91%73   </td>74  </tr>75</table>76 77 78### Reproduction79 80The GSM8K results were obtained using the vLLM framework, based on the Docker image `rocm/vllm-dev:nightly_main_20260211`, and vLLM is installed inside the container.81 82#### Preparation in container83 84To download the evaluation script, reinstallation is not required.85```86# Install vLLM code repo87git clone https://github.com/vllm-project/vllm.git88cd vllm89git checkout v0.13.090cd ..91```92 93#### Launching server94```95VLLM_ROCM_USE_AITER=1 \96VLLM_DISABLE_COMPILE_CACHE=1 \97vllm serve "$MODEL" \98    --tensor-parallel-size 4 \99    --trust-remote-code \100    --max-model-len 32768 \101    --port 8899  102```103 104 105#### Evaluating model in a new terminal106```107python vllm/tests/evals/gsm8k/gsm8k_eval.py --host http://127.0.0.1 --port 8899 --num-questions 1000 --save-results logs108```109 110 111# License112Modifications Copyright(c) 2026 Advanced Micro Devices, Inc. All rights reserved.