Ferox-AI/Phi-4-multimodal-instruct-mlx-4bit
Phi-4-Multimodal-Instruct — MLX 4-bit
A 4-bit quantized Apple MLX conversion of microsoft/Phi-4-multimodal-instruct for native inference on Apple Silicon.
Converted by [Ferox AI](https://ferox.ca) · Vision-language inference on MacBook / Mac Studio / Mac Pro without cloud dependencies.
Other variants: bf16 (full precision) · 8-bit
Quickstart
from mlx_vlm import load, generate
model, processor = load("ferox-ai/Phi-4-multimodal-instruct-mlx-4bit")
output = generate(
model,
processor,
"Describe this image in detail.",
["path/to/image.jpg"],
max_tokens=512,
verbose=False,
)
print(output)Requires mlx-vlm >= 0.1.0 with Phi-4-MM architecture support. Install dependencies:
pip install mlx-vlm>=0.1.0 mlx>=0.22.0Benchmark Results
Evaluated with our internal evaluation harness on a single Apple Silicon device. Scores are computed on a 100-sample subset of each benchmark. Microsoft's reference scores are reported on the full dataset using PyTorch FP16 — direct comparison should account for both the precision difference and sample-size variance.
† ScienceQA: 48 of 100 samples scored (image-bearing questions only; 52 text-only questions excluded).
Quantization impact
Across all benchmarks, 4-bit quantization produces a mean accuracy delta of −2.2 percentage points relative to bf16 — within the expected range for 4-bit group quantization on a model of this scale.
Note on MMMU
The 100-sample MMMU scores (24.0% 4-bit, 31.0% bf16) fall well below Microsoft's reported 55.1%. To isolate the cause, we ran a full 900-sample MMMU validation on the lossless bf16 variant and obtained 27.9% — consistent with the subset, which confirms the gap is not caused by quantization or weight conversion. We were unable to reproduce Microsoft's 55.1% and attribute the difference to evaluation-harness and answer-extraction handling for MMMU's multiple-choice format (prompt formatting and option parsing), rather than to the model's underlying capability — which is better reflected by the document-, chart-, OCR-, and science-focused benchmarks above.
Architecture
Weight provenance
Weights are converted from microsoft/Phi-4-multimodal-instruct using a deterministic pipeline:
- Download source checkpoint (PyTorch safetensors)
- Fuse vision LoRA adapters into backbone weights (eliminates runtime adapter overhead)
- Remap weight keys to MLX naming conventions
- Transpose LoRA matrices (PEFT → MLX format)
- Quantize backbone to 4-bit (SigLIP excluded)
- Serialize as MLX safetensors
The conversion and quantization pipeline is deterministic and fully reproducible from the base model.
Intended Use
This model is designed for local, on-device vision-language inference on Apple Silicon hardware. Suitable applications include:
- Document understanding and extraction (invoices, forms, reports)
- Chart and diagram interpretation
- Visual question answering
- OCR and text recognition in images
- Educational content analysis
Out of scope
- Audio processing (Phase 2, not included in this release)
- Production deployment without application-level safety filtering
- Use cases requiring guaranteed factual accuracy without human verification
Limitations
- 100-sample evaluations. Benchmark scores are computed on subsets, not full datasets. Expect variance relative to full-dataset evaluations.
- Vision-only. This is a Phase 1 release covering the vision modality. Audio support from the original Phi-4-multimodal architecture is not included.
- No runtime LoRA switching. Vision LoRA adapters are pre-fused; the model cannot dynamically swap adapters.
- Apple Silicon required. MLX is designed for Apple's unified memory architecture (M1/M2/M3/M4). This model will not run on CUDA or CPU-only systems.
Citation
If you use this model in your work, please cite:
@misc{feroxai2026phi4mlx,
title={Phi-4-Multimodal-Instruct MLX Conversion},
author={Ferox AI},
year={2026},
url={https://huggingface.co/ferox-ai/Phi-4-multimodal-instruct-mlx-4bit},
note={4-bit quantized MLX port of microsoft/Phi-4-multimodal-instruct}
}Acknowledgments
- Microsoft Research for the Phi-4-multimodal-instruct model and technical report
- Apple MLX team for the MLX framework
- Prince Canuma for mlx-vlm
