Irfanuruchi/Polaris-V1-MLX-8bit
Polaris-V1 MLX 8-bit
This is an MLX 8-bit conversion of `nitrai-research/Polaris-V1` for inference on Apple silicon.
Quantization
- Format: MLX
- Quantization mode: affine
- Bits: 8
- Group size: 64
- Effective size: 8.502 bits per weight
- Output size: approximately 4.2 GB
Compatibility repairs
The source checkpoint identifies its text architecture as qwen3_5_text. MLX-LM exposes the compatible text implementation as qwen3_5, so the converted configuration was normalized accordingly.
The source configuration also used <|endoftext|> (248044) as EOS while the tokenizer and chat template terminate responses with <|im_end|> (248046). Both config.json and generation_config.json were corrected to use 248046, preventing repeated stop-token generation.
Installation
~~~bash pip install -U mlx-lm ~~~
Python usage
~~~python from mlx_lm import load, generate
model, tokenizer = load("Irfanuruchi/Polaris-V1-MLX-8bit")
response = generate( model, tokenizer, prompt="Explain virtual memory in three concise points.", max_tokens=256, )
print(response) ~~~
Command-line usage
~~~bash mlx_lm.generate \ --model Irfanuruchi/Polaris-V1-MLX-8bit \ --prompt "Explain virtual memory in three concise points." \ --max-tokens 256 ~~~
Validation
Validated locally on an Apple M3 Pro MacBook Pro using:
- Python 3.12.14
- MLX 0.32.1
- MLX-LM 0.31.3
- Generation speed: 27.44 tokens/second
- Peak unified memory: 4.63 GB
- EOS termination: passed
- Coherent text generation: passed
Benchmark results in the metadata above are inherited from the source model card and were not independently reproduced for this quantized conversion.
License
Apache 2.0, inherited from the source model.
