CoolFace
Modelpublic

EnlistedGhost/Ministral-3-3B-Instruct-2512-GGUF

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes1kdownloads
Model Card

<img src="https://ollama.com/assets/library/mistral-nemo/72045292-694a-4867-88c8-8635c9d97030" alt="Example image" width="168" height="128">

<img src="https://ollama.com/assets/library/ministral-3/83fa3859-d87f-492c-bd81-596cfbceeccb" alt="Example image" width="64" height="64">

-----------------------------------------------<br /> - Update August 14th 2026 -<br />-----------------------------------------------

  • —New chat template! Yes, a new chat template that completely fixes previous issues for llama.cpp users <br />and retains higher performance than the originally made template for Ollama/Yollama users!
  • —Fixed both F16 and F32 MMPROJ Multi-Modal Vision projectors. (Re-Uploaded with working projectors)
  • —Uploaded custom, high-quality, new "GH05T" edition Quant files for this model! (This is a brand new method I have been developing for some time, I sincerely hope you enjoy it!)

Custom GH05T Quants Added: These offer significantly higher performance and quality of response over previously uploaded release files while remaining <br /> similar or smaller in size!

  • —IQ2_M
  • —Q2KL
  • —IQ3_M
  • —Q3KL
  • —IQ4_XS
  • —Q4KM
  • —Q5KXL
  • —Q6KL
  • —Q80L
  • —F16

New JINJA Tokenizer Chat-Template: <br /> (This template features a sliding context window of TWENTY-NINE (29) messages. <br /> This can be adjusted per-individual requirements simply by altering <br />the number 29 in the template higher or lower in numerical value)

jinja
{%- set ns = namespace(remMessage=false, hasSys=false, injSystem=true) -%}
{%- for msg in messages -%}
	{%- if msg.role == "system" -%}
		{%- set ns.hasSys = true -%}
	{%- endif -%}
{%- endfor -%}
{%- for msg in messages -%}
	{{- bos_token }}
	{%- if ns.injSystem -%}
		[SYSTEM_PROMPT]
		{%- if ns.hasSys -%}
			{{ msg.content }}
		{%- else -%}
			'Follow instructions the user provides... (System Prompt)'
		{%- endif -%}
		[/SYSTEM_PROMPT]
		{%- set ns.injSystem = false -%}
	{%- endif -%}
	{%- if (messages|length - loop.index0) < 29 -%}
		{%- set ns.remMessage = true -%}
	{%- endif -%}
	{%- if ns.remMessage -%}
		{%- if msg.role == "user" -%}
			{{- '[INST]' }}
			{%- if msg.content is string %}
				{{ msg.content }}
        	{%- else %}
            	{%- for block in msg.content %}
            		{%- if block.type == 'text' %}
                    	{{- block.text }}
                	{%- elif block.type in ['image', 'image_url'] %}
                    	{{- '[IMG]' }}
                	{%- endif %}
            	{%- endfor %}
            {%- endif %}
            {{- '[/INST]' }}
		{%- elif msg.role == "assistant" -%}
			{{ msg.content }}
		{%- endif -%}
	{%- endif -%}
	{{- eos_token }}
{%- endfor -%}

------------------------------------------------<br /> - Model Details and Specifications: -<br />------------------------------------------------

Ministral-3 3B Instruct 2512 (GGUF)

This release contains: <br /> GGUF converted and Quantized model files (Compatible with:)

Quantized GGUF version of:

  • —Ministral-3-3B-Instruct-2512-BF16 <br /> (by MistralAI)

Original Model Link:


-------------------------------------------------------------<br /> - GGUF Conversion and Quantization Details: -<br />-------------------------------------------------------------

Software used to convert Safetensors to GGUF:

  • —<a href="https://github.com/ggml-org/llama.cpp/">llama.cpp</a>

Software used to create Quantized GGUF Files:

  • —<a href="https://github.com/ggml-org/llama.cpp/">llama.cpp</a>

Specific GitHub Commit Point:

  • —<a href="https://github.com/ggml-org/llama.cpp/commit/85c40c9b02941ebf1add1469af75f1796d513ef4">b7540</a>

Converted to GGUF and Quantized by:


--------------------------<br /> ---- Original Info ---- <br /> --------------------------

(Crossposted from the link in the above section: "Model Details"): <br /> <br /> <br /> <br />

Ministral 3 14B Instruct 2512 BF16

The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language model with vision capabilities.

This model is the instruct post-trained version, fine-tuned for instruction tasks, making it ideal for chat and instruction based use cases.

The Ministral 3 family is designed for edge deployment, capable of running on a wide range of hardware. Ministral 3 14B can even be deployed locally, capable of fitting in 32GB of VRAM in BF16, and less than 24GB of RAM/VRAM when quantized.

We provide a no-loss FP8 version here, you can find other formats and quantizations in the Ministral 3 - Additional Checkpoints collection.

Key Features

Ministral 3 14B consists of two main architectural components:

  • —13.5B Language Model
  • —0.4B Vision Encoder

The Ministral 3 14B Instruct model offers the following capabilities:

  • —Vision: Enables the model to analyze images and provide insights based on visual content, in addition to text.
  • —Multilingual: Supports dozens of languages, including English, French, Spanish, German, Italian, Portuguese, Dutch, Chinese, Japanese, Korean, Arabic.
  • —System Prompt: Maintains strong adherence and support for system prompts.
  • —Agentic: Offers best-in-class agentic capabilities with native function calling and JSON outputting.
  • —Edge-Optimized: Delivers best-in-class performance at a small scale, deployable anywhere.
  • —Apache 2.0 License: Open-source license allowing usage and modification for both commercial and non-commercial purposes.
  • —Large Context Window: Supports a 256k context window.

Use Cases

Private AI deployments where advanced capabilities meet practical hardware constraints:

  • —Private/custom chat and AI assistant deployments in constrained environments
  • —Advanced local agentic use cases
  • —Fine-tuning and specialization
  • —And more...

Bringing advanced AI capabilities to most environments.

Ministral 3 Family

Model NameTypePrecisionLink
Ministral 3 3B Base 2512Base pre-trainedBF16Hugging Face
Ministral 3 3B Instruct 2512Instruct post-trainedBF16Hugging Face
Ministral 3 3B Reasoning 2512Reasoning capableBF16Hugging Face
Ministral 3 8B Base 2512Base pre-trainedBF16Hugging Face
Ministral 3 8B Instruct 2512Instruct post-trainedBF16Hugging Face
Ministral 3 8B Reasoning 2512Reasoning capableBF16Hugging Face
Ministral 3 14B Base 2512Base pre-trainedBF16Hugging Face
Ministral 3 14B Instruct 2512Instruct post-trainedBF16Hugging Face
Ministral 3 14B Reasoning 2512Reasoning capableBF16Hugging Face

Other formats available here.

Benchmark Results

We compare Ministral 3 to similar sized models.

Reasoning

ModelAIME25AIME24GPQA DiamondLiveCodeBench
Ministral 3 14B<u>0.850</u><u>0.898</u><u>0.712</u><u>0.646</u>
Qwen3-14B (Thinking)0.7370.8370.6630.593
Ministral 3 8B0.787<u>0.860</u>0.668<u>0.616</u>
Qwen3-VL-8B-Thinking<u>0.798</u><u>0.860</u><u>0.671</u>0.580
Ministral 3 3B<u>0.721</u><u>0.775</u>0.534<u>0.548</u>
Qwen3-VL-4B-Thinking0.6970.729<u>0.601</u>0.513

Instruct

ModelArena HardWildBenchMATH Maj@1MM MTBench
Ministral 3 14B<u>0.551</u><u>68.5</u><u>0.904</u><u>8.49</u>
Qwen3 14B (Non-Thinking)0.42765.10.870NOT MULTIMODAL
Gemma3-12B-Instruct0.43663.20.8546.70
Ministral 3 8B0.509<u>66.8</u>0.876<u>8.08</u>
Qwen3-VL-8B-Instruct<u>0.528</u>66.3<u>0.946</u>8.00
Ministral 3 3B0.305<u>56.8</u>0.8307.83
Qwen3-VL-4B-Instruct<u>0.438</u><u>56.8</u><u>0.900</u><u>8.01</u>
Qwen3-VL-2B-Instruct0.16342.20.7866.36
Gemma3-4B-Instruct0.31849.10.7595.23

Base

ModelMultilingual MMLUMATH CoT 2-ShotAGIEval 5-shotMMLU Redux 5-shotMMLU 5-shotTriviaQA 5-shot
Ministral 3 14B0.742<u>0.676</u>0.6480.8200.7940.749
Qwen3 14B Base<u>0.754</u>0.620<u>0.661</u><u>0.837</u><u>0.804</u>0.703
Gemma 3 12B Base0.6900.4870.5870.7660.745<u>0.788</u>
Ministral 3 8B<u>0.706</u><u>0.626</u>0.5910.793<u>0.761</u><u>0.681</u>
Qwen 3 8B Base0.7000.576<u>0.596</u><u>0.794</u>0.7600.639
Ministral 3 3B0.652<u>0.601</u>0.5110.7350.7070.592
Qwen 3 4B Base<u>0.677</u>0.405<u>0.570</u><u>0.759</u><u>0.713</u>0.530
Gemma 3 4B Base0.5160.2940.4300.6260.589<u>0.640</u>

License

This model is licensed under the Apache 2.0 License.

You must not use this model in a manner that infringes, misappropriates, or otherwise violates any third party’s rights, including intellectual property rights.