CoolFace
Modelpublic

mlx-community/Fara1.5-27B-OptiQ-4bit

sourceHugging Facemitupdated 10d agoView on Hugging Face
2likes322downloads
Model Card

Fara1.5-27B-OptiQ-4bit

Built with [mlx-optiq](https://mlx-optiq.com), the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon (no PyTorch, no cloud). Try the Lab · All OptiQ quants · Docs Supported loaders: mlx-optiq (text, vision, and MTP) and stock mlx-lm (text). Other front-ends load MLX weights through their own stack, so support there depends on that stack rather than on these files.

An OptiQ mixed-precision MLX quant of microsoft/Fara1.5-27B, a Qwen3.5-based computer-use / web-agent vision-language model.

  • —Mixed 4/8-bit, 4.74 bits per weight (20G on disk).
  • —The per-layer bit allocation is transferred from the published `mlx-community/Qwen3.5-27B-OptiQ-4bit` quant. Fara1.5 is a finetune of Qwen3.5-27B with identical architecture, so the OptiQ allocation matches the Qwen3.5 family exactly, with no separate sensitivity pass.
  • —Vision tower kept at bf16 in optiq/optiq_vision.safetensors. The one repo loads text-only under stock mlx-lm and full image+text under OptiQ.

Running it

bash
pip install -U optiq
optiq serve --model mlx-community/Fara1.5-27B-OptiQ-4bit

Use the OpenAI-compatible endpoint at http://localhost:8000/v1. Send an image_url part for the computer-use / vision path.