mlx-community/Ministral-3-3B-Base-2512-4bit
0125
mlx-community/Ministral-3-3B-Base-2512-4bit
This is, Ministral 3 3B Base 2512 is a vision-language model: a text backbone paired with a vision encoder, supporting image understanding alongside text. This is the base pre-trained checkpoint — not instruction- or chat-tuned. For chat/instruction-following use cases, use the Instruct variant instead; this base checkpoint is intended for custom post-training/fine-tuning.
Community note. Structural check confirms the vision tower and multimodal projector were carried over intact (not dropped, which is a real failure mode for text-only conversion tools on vision-language models). Functional check confirms both text-only and image+text generation produce coherent output. Converted and verified by a single maintainer running local MLX tooling -- not independently reviewed by anyone else; please open a discussion if you hit anything unexpected.
This is an MLX conversion of `mistralai/Ministral-3-3B-Base-2512`, converted with mlx-vlm. Refer to the original model card for the full description, capabilities, and license terms.
Heads up
- Base model, not instruct-tuned — expect raw completion behavior, not chat-following. Don't expect it to follow instructions well.
- Vision retained at full precision — only the language backbone is quantized; the vision tower and multimodal projector are untouched bf16, per mlx-vlm's standard policy of not quantizing multimodal modules.
- Output size on disk: 2.80GB
Provenance
- Source: `mistralai/Ministral-3-3B-Base-2512` (BF16)
- Language model layers: 4-bit affine quantization, group_size=64
- Vision tower + multimodal projector: kept at full precision (not quantized)
- Blended average: 5.756 bits per weight across all parameters
Ministral 3 family
Use with mlx
pip install -U mlx-vlmpython -m mlx_vlm.generate --model mlx-community/Ministral-3-3B-Base-2512-4bit --max-tokens 100 --temperature 0.0 --prompt "Describe this image." --image <path_to_image>For text-only prompts, omit --image.
License
Apache 2.0 — see the original model card for the full license text and any usage terms.
