mlx-community/FastContext-1.0-4B-SFT-8bit
FastContext-1.0-4B-SFT — MLX 8-bit
MLX 8-bit quantization (8.500 bits per weight, ~4 GB) of ShaunGves/FastContext-1.0-4B-SFT, a surviving mirror of Microsoft's FastContext repository-exploration subagent (arXiv:2606.14066), removed from the official listings on 2026-06-30. Base model: Qwen3-4B-Instruct-2507. Converted with mlx-lm 0.29.1 (default group size 64).
FastContext is a small explorer model for coding agents: given a natural-language query about a repository, it explores with read-only tools (Read / Glob / Grep, called in parallel) and returns a compact <final_answer> block of file:line-range citations, keeping broad exploration out of the main agent's context window.
Quality note (tested)
Verified end-to-end with the fastcontext CLI on an M4 Mac: at temperature 0.6 this 8-bit quant produced accurate, verifiable citations on par with the FP16 weights (~15 s per query, 4 exploration turns). This is the recommended quant. The 4-bit sibling showed degraded path grounding in the same tests — prefer 8-bit unless memory is tight.
Use with LM Studio
Search for FastContext-1.0-4B-SFT-8bit in LM Studio (MLX runtime), or:
lms get mlx-community/FastContext-1.0-4B-SFT-8bitUse with mlx-lm
uv tool install "mlx-lm==0.29.1" --with "transformers<5" --with "mlx<0.31"
mlx_lm.server --model mlx-community/FastContext-1.0-4B-SFT-8bit --port 8080Use with the fastcontext CLI
Install from the preserved mirror (includes fixes for local OpenAI-compatible servers), then from the repo you want to explore:
export BASE_URL="http://localhost:8080/v1" # or http://localhost:1234/v1 for LM Studio
export MODEL="mlx-community/FastContext-1.0-4B-SFT-8bit"
export API_KEY="local"
export TEMPERATURE=0.6
export MAX_TOKENS=4000
fastcontext -q "Where is the retry logic for failed API calls?" --citation