iapp/openthai2.0-qwen3.8-27b-MLX-4bit
0247
OpenThai2.0 - Opensource Thai Knowledge, Document, and Agentic AI (MLX 4-bit)
MLX 4-bit quantization (mlx-vlm) of iapp/openthai2.0-qwen3.8-27b for Apple silicon. Vision included. ~16 GB — runs on 24 GB+ unified memory.
v2.0.3 — rebuilt from the v2.0.3 weights (Thai knowledge-recall fix). The v2.0.0 launch build stays available at revision tag v2.0.0.Run
pip install mlx-vlm python -m mlx_vlm generate --model iapp/openthai2.0-qwen3.8-27b-MLX-4bit \ --image document.jpg --prompt "อ่านข้อความในเอกสารนี้ทั้งหมด" --max-tokens 8192
⚠️ The model reasons before it answers — leave a large generation budget (8k+), or replies may come back empty.
Notes: the MTP draft head is not included (mlx-vlm has no drafter support for this architecture yet). Sanity-verified on-device: Thai factual prompts answered correctly. Full benchmarks and model card: see the main bf16 repo.
