CoolFace
Modelpublic

dbaek111/fastvlm-0.5b-mlx-q8

sourceHugging Faceapple-amlrupdated 2mo agoView on Hugging Face
0likes16downloads
Model Card

fastvlm-0.5b-mlx-q8

This repository contains an MLX-converted FastVLM checkpoint.

Model

  • —Base model: apple/FastVLM-0.5B
  • —Parameters: 0.5B
  • —Precision: q8 / 8-bit quantized
  • —Approx. folder size: 1.1G

The checkpoint was converted from Apple FastVLM using the official FastVLM model export workflow and patched mlx-vlm.

Files

This repository should include:

  • —config.json
  • —MLX model weights
  • —tokenizer files
  • —fastvithd.mlpackage vision tower

Example Usage

bash
hf download dbaek111/fastvlm-0.5b-mlx-q8 --local-dir ./fastvlm-0.5b-mlx-q8

python -m mlx_vlm.generate \
  --model ./fastvlm-0.5b-mlx-q8 \
  --image /path/to/your/image.jpg \
  --prompt "Explain the image." \
  --max-tokens 64 \
  --temp 0.0

Benchmark

Benchmarked on an Apple Silicon Mac with the patched FastVLM mlx-vlm workflow.

  • —Task: pedestrian wayfinding captioning
  • —Images: three local test images resized to 512px and 1024px long edge
  • —Prompt: Describe what is visible for pedestrian wayfinding in one short sentence. Do not list categories. Do not mention anything you cannot see. Keep under 30 words.
  • —Max tokens: 64
  • —Temperature: 0.0
  • —Timing: model loaded once per image set, then three images processed sequentially

Model Selection

ModelSizePrecisionAvg 512pxAvg 1024pxLoadRecommended use
fastvlm-0.5b-mlx-q4819Mq40.338s0.370s2.82sSmallest and fastest; rough real-time captions
fastvlm-0.5b-mlx-q81.1Gq80.419s0.414s2.57sFast, with richer captions than 0.5B q4
fastvlm-0.5b-mlx-fp161.6Gfp160.435s0.421s2.79sSmall FP16 baseline
fastvlm-1.5b-mlx-q41.4Gq40.447s0.464s2.65sBest real-time balance for pedestrian wayfinding
fastvlm-1.5b-mlx-q82.2Gq80.552s0.541s2.68sMore detail while staying sub-second
fastvlm-1.5b-mlx-fp163.8Gfp160.636s0.557s3.08s1.5B FP16 reference variant
fastvlm-7b-mlx-q44.9Gq41.263s1.241s3.04sBest quality/latency tradeoff among 7B variants
fastvlm-7b-mlx-q88.0Gq81.497s1.495s3.85sHigher precision 7B, slower than q4
fastvlm-7b-mlx-fp1615Gfp161.834s1.874s45.48sFull precision reference; expensive to load

Per-Image Timing

Each cell is img1 / img2 / img3 inference time in seconds.

Model512px images1024px images
fastvlm-0.5b-mlx-q40.318 / 0.353 / 0.3430.374 / 0.370 / 0.366
fastvlm-0.5b-mlx-q80.393 / 0.465 / 0.4000.424 / 0.387 / 0.431
fastvlm-0.5b-mlx-fp160.394 / 0.490 / 0.4210.430 / 0.395 / 0.437
fastvlm-1.5b-mlx-q40.454 / 0.458 / 0.4300.465 / 0.472 / 0.456
fastvlm-1.5b-mlx-q80.591 / 0.611 / 0.4530.594 / 0.552 / 0.477
fastvlm-1.5b-mlx-fp160.701 / 0.709 / 0.4960.520 / 0.635 / 0.516
fastvlm-7b-mlx-q41.161 / 1.330 / 1.2971.081 / 1.343 / 1.298
fastvlm-7b-mlx-q81.364 / 1.639 / 1.4871.341 / 1.712 / 1.432
fastvlm-7b-mlx-fp161.561 / 2.041 / 1.9021.595 / 2.277 / 1.749

Compatibility

This is an MLX export of FastVLM for Apple Silicon Macs. It includes the CoreML FastViTHD vision tower as fastvithd.mlpackage.

This repository is not a standard PyTorch Transformers checkpoint and is not intended for vLLM, SGLang, or Linux GPU inference.

Notes

This is a converted and quantized derivative of Apple FastVLM.

Please refer to the original Apple FastVLM repository and model card for license and usage conditions.