CoolFace
Modelpublic

cyankiwi/Ministral-3-3B-Instruct-2512-AWQ-4bit

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes442downloads
Model Card

Ministral-3-3B-Instruct-2512 AWQ - INT4

Model Details

Quantization Details

Memory Usage

**Type****Ministral-3-3B-Instruct-2512-BF16****Ministral-3-3B-Instruct-2512-AWQ-4bit**
Memory Size14.3 GB7.7 GB

Evaluations

**Benchmarks****Ministral-3-3B-Instruct-2512-BF16****Ministral-3-3B-Instruct-2512-AWQ-4bit**
Perplexity1.587461.59943
  • —Evaluation Context Length: 16384

Inference

Prerequisite

bash
pip install -U vllm

Basic Usage

bash
vllm serve cyankiwi/Ministral-3-3B-Instruct-2512-AWQ-4bit --tokenizer_mode mistral --config_format mistral --load_format mistral --enable-auto-tool-choice --tool-call-parser mistral

Additional Information

Changelog

  • —v1.0.0 - Initial quantized release

Authors

  • —Name: Ton Cao
  • —Contacts: ton@cyan.kiwi

Ministral 3 3B Instruct 2512 BF16

The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.

This model is the instruct post-trained version, fine-tuned for instruction tasks, making it ideal for chat and instruction based use cases.

The Ministral 3 family is designed for edge deployment, capable of running on a wide range of hardware. Ministral 3 3B can even be deployed locally, capable of fitting in 16GB of VRAM in BF16, and less than 8GB of RAM/VRAM when quantized.

We provide a no-loss FP8 version here, you can find other formats and quantizations in the Ministral 3 - Additional Checkpoints collection.

Key Features

Ministral 3 3B consists of two main architectural components:

  • —3.4B Language Model
  • —0.4B Vision Encoder

The Ministral 3 3B Instruct model offers the following capabilities:

  • —Vision: Enables the model to analyze images and provide insights based on visual content, in addition to text.
  • —Multilingual: Supports dozens of languages, including English, French, Spanish, German, Italian, Portuguese, Dutch, Chinese, Japanese, Korean, Arabic.
  • —System Prompt: Maintains strong adherence and support for system prompts.
  • —Agentic: Offers best-in-class agentic capabilities with native function calling and JSON outputting.
  • —Edge-Optimized: Delivers best-in-class performance at a small scale, deployable anywhere.
  • —Apache 2.0 License: Open-source license allowing usage and modification for both commercial and non-commercial purposes.
  • —Large Context Window: Supports a 256k context window.

Use Cases

Ideal for lightweight, real-time applications on edge or low-resource devices, such as:

  • —Image captioning
  • —Text classification
  • —Real-time efficient translation
  • —Data extraction
  • —Short content generation
  • —Fine-tuning and specialization
  • —And more...

Bringing advanced AI capabilities to edge and distributed environments for embedded systems.

Ministral 3 Family

Model NameTypePrecisionLink
Ministral 3 3B Base 2512Base pre-trainedBF16Hugging Face
Ministral 3 3B Instruct 2512Instruct post-trainedBF16Hugging Face
Ministral 3 3B Reasoning 2512Reasoning capableBF16Hugging Face
Ministral 3 8B Base 2512Base pre-trainedBF16Hugging Face
Ministral 3 8B Instruct 2512Instruct post-trainedBF16Hugging Face
Ministral 3 8B Reasoning 2512Reasoning capableBF16Hugging Face
Ministral 3 14B Base 2512Base pre-trainedBF16Hugging Face
Ministral 3 14B Instruct 2512Instruct post-trainedBF16Hugging Face
Ministral 3 14B Reasoning 2512Reasoning capableBF16Hugging Face

Other formats available here.

Benchmark Results

We compare Ministral 3 to similar sized models.

Reasoning

ModelAIME25AIME24GPQA DiamondLiveCodeBench
Ministral 3 14B<u>0.850</u><u>0.898</u><u>0.712</u><u>0.646</u>
Qwen3-14B (Thinking)0.7370.8370.6630.593
Ministral 3 8B0.787<u>0.860</u>0.668<u>0.616</u>
Qwen3-VL-8B-Thinking<u>0.798</u><u>0.860</u><u>0.671</u>0.580
Ministral 3 3B<u>0.721</u><u>0.775</u>0.534<u>0.548</u>
Qwen3-VL-4B-Thinking0.6970.729<u>0.601</u>0.513

Instruct

ModelArena HardWildBenchMATH Maj@1MM MTBench
Ministral 3 14B<u>0.551</u><u>68.5</u><u>0.904</u><u>8.49</u>
Qwen3 14B (Non-Thinking)0.42765.10.870NOT MULTIMODAL
Gemma3-12B-Instruct0.43663.20.8546.70
Ministral 3 8B0.509<u>66.8</u>0.876<u>8.08</u>
Qwen3-VL-8B-Instruct<u>0.528</u>66.3<u>0.946</u>8.00
Ministral 3 3B0.305<u>56.8</u>0.8307.83
Qwen3-VL-4B-Instruct<u>0.438</u><u>56.8</u><u>0.900</u><u>8.01</u>
Qwen3-VL-2B-Instruct0.16342.20.7866.36
Gemma3-4B-Instruct0.31849.10.7595.23

Base

ModelMultilingual MMLUMATH CoT 2-ShotAGIEval 5-shotMMLU Redux 5-shotMMLU 5-shotTriviaQA 5-shot
Ministral 3 14B0.742<u>0.676</u>0.6480.8200.7940.749
Qwen3 14B Base<u>0.754</u>0.620<u>0.661</u><u>0.837</u><u>0.804</u>0.703
Gemma 3 12B Base0.6900.4870.5870.7660.745<u>0.788</u>
Ministral 3 8B<u>0.706</u><u>0.626</u>0.5910.793<u>0.761</u><u>0.681</u>
Qwen 3 8B Base0.7000.576<u>0.596</u><u>0.794</u>0.7600.639
Ministral 3 3B0.652<u>0.601</u>0.5110.7350.7070.592
Qwen 3 4B Base<u>0.677</u>0.405<u>0.570</u><u>0.759</u><u>0.713</u>0.530
Gemma 3 4B Base0.5160.2940.4300.6260.589<u>0.640</u>

License

This model is licensed under the Apache 2.0 License.

You must not use this model in a manner that infringes, misappropriates, or otherwise violates any third party’s rights, including intellectual property rights.