CoolFace
Modelpublic

0utsideness/Ministral-3-3B-Instruct-2512-BF16-heretic-experimental

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
1likes9downloads
Model Card

This is a decensored version of mistralai/Ministral-3-3B-Instruct-2512-BF16, made using a development version of Heretic

Abliteration parameters

ParameterValue
direction_index13.22
attn.o_proj.max_weight1.13
attn.o_proj.max_weight_position15.64
attn.o_proj.min_weight1.11
attn.o_proj.min_weight_distance14.19
mlp.down_proj.max_weight0.99
mlp.down_proj.max_weight_position15.31
mlp.down_proj.min_weight0.52
mlp.down_proj.min_weight_distance13.77

Performance

MetricThis modelOriginal model ([mistralai/Ministral-3-3B-Instruct-2512-BF16](https://huggingface.co/mistralai/Ministral-3-3B-Instruct-2512-BF16))
KL divergence0.01540 (by definition)
Refusals28/10099/100

Ministral 3 3B Instruct 2512 BF16

The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.

This model is the instruct post-trained version, fine-tuned for instruction tasks, making it ideal for chat and instruction based use cases.

The Ministral 3 family is designed for edge deployment, capable of running on a wide range of hardware. Ministral 3 3B can even be deployed locally, capable of fitting in 16GB of VRAM in BF16, and less than 8GB of RAM/VRAM when quantized.

We provide a no-loss FP8 version here, you can find other formats and quantizations in the Ministral 3 - Additional Checkpoints collection.

Learn more in our blog post and paper.

Key Features

Ministral 3 3B consists of two main architectural components:

  • —3.4B Language Model
  • —0.4B Vision Encoder

The Ministral 3 3B Instruct model offers the following capabilities:

  • —Vision: Enables the model to analyze images and provide insights based on visual content, in addition to text.
  • —Multilingual: Supports dozens of languages, including English, French, Spanish, German, Italian, Portuguese, Dutch, Chinese, Japanese, Korean, Arabic.
  • —System Prompt: Maintains strong adherence and support for system prompts.
  • —Agentic: Offers best-in-class agentic capabilities with native function calling and JSON outputting.
  • —Edge-Optimized: Delivers best-in-class performance at a small scale, deployable anywhere.
  • —Apache 2.0 License: Open-source license allowing usage and modification for both commercial and non-commercial purposes.
  • —Large Context Window: Supports a 256k context window.

Use Cases

Ideal for lightweight, real-time applications on edge or low-resource devices, such as:

  • —Image captioning
  • —Text classification
  • —Real-time efficient translation
  • —Data extraction
  • —Short content generation
  • —Fine-tuning and specialization
  • —And more...

Bringing advanced AI capabilities to edge and distributed environments for embedded systems.

Ministral 3 Family

Model NameTypePrecisionLink
Ministral 3 3B Base 2512Base pre-trainedBF16Hugging Face
Ministral 3 3B Instruct 2512Instruct post-trainedBF16Hugging Face
Ministral 3 3B Reasoning 2512Reasoning capableBF16Hugging Face
Ministral 3 8B Base 2512Base pre-trainedBF16Hugging Face
Ministral 3 8B Instruct 2512Instruct post-trainedBF16Hugging Face
Ministral 3 8B Reasoning 2512Reasoning capableBF16Hugging Face
Ministral 3 14B Base 2512Base pre-trainedBF16Hugging Face
Ministral 3 14B Instruct 2512Instruct post-trainedBF16Hugging Face
Ministral 3 14B Reasoning 2512Reasoning capableBF16Hugging Face

Other formats available here.

Benchmark Results

We compare Ministral 3 to similar sized models.

Reasoning

ModelAIME25AIME24GPQA DiamondLiveCodeBench
Ministral 3 14B<u>0.850</u><u>0.898</u><u>0.712</u><u>0.646</u>
Qwen3-14B (Thinking)0.7370.8370.6630.593
Ministral 3 8B0.787<u>0.860</u>0.668<u>0.616</u>
Qwen3-VL-8B-Thinking<u>0.798</u><u>0.860</u><u>0.671</u>0.580
Ministral 3 3B<u>0.721</u><u>0.775</u>0.534<u>0.548</u>
Qwen3-VL-4B-Thinking0.6970.729<u>0.601</u>0.513

Instruct

ModelArena HardWildBenchMATH Maj@1MM MTBench
Ministral 3 14B<u>0.551</u><u>68.5</u><u>0.904</u><u>8.49</u>
Qwen3 14B (Non-Thinking)0.42765.10.870NOT MULTIMODAL
Gemma3-12B-Instruct0.43663.20.8546.70
Ministral 3 8B0.509<u>66.8</u>0.876<u>8.08</u>
Qwen3-VL-8B-Instruct<u>0.528</u>66.3<u>0.946</u>8.00
Ministral 3 3B0.305<u>56.8</u>0.8307.83
Qwen3-VL-4B-Instruct<u>0.438</u><u>56.8</u><u>0.900</u><u>8.01</u>
Qwen3-VL-2B-Instruct0.16342.20.7866.36
Gemma3-4B-Instruct0.31849.10.7595.23

Base

ModelMultilingual MMLUMATH CoT 2-ShotAGIEval 5-shotMMLU Redux 5-shotMMLU 5-shotTriviaQA 5-shot
Ministral 3 14B0.742<u>0.676</u>0.6480.8200.7940.749
Qwen3 14B Base<u>0.754</u>0.620<u>0.661</u><u>0.837</u><u>0.804</u>0.703
Gemma 3 12B Base0.6900.4870.5870.7660.745<u>0.788</u>
Ministral 3 8B<u>0.706</u><u>0.626</u>0.5910.793<u>0.761</u><u>0.681</u>
Qwen 3 8B Base0.7000.576<u>0.596</u><u>0.794</u>0.7600.639
Ministral 3 3B0.652<u>0.601</u>0.5110.7350.7070.592
Qwen 3 4B Base<u>0.677</u>0.405<u>0.570</u><u>0.759</u><u>0.713</u>0.530
Gemma 3 4B Base0.5160.2940.4300.6260.589<u>0.640</u>

License

This model is licensed under the Apache 2.0 License.

You must not use this model in a manner that infringes, misappropriates, or otherwise violates any third party’s rights, including intellectual property rights.