CoolFace
Modelpublic

Irfanuruchi/Polaris-V1-MLX-8bit

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes11downloads
Model Card

Polaris-V1 MLX 8-bit

This is an MLX 8-bit conversion of `nitrai-research/Polaris-V1` for inference on Apple silicon.

Quantization

  • —Format: MLX
  • —Quantization mode: affine
  • —Bits: 8
  • —Group size: 64
  • —Effective size: 8.502 bits per weight
  • —Output size: approximately 4.2 GB

Compatibility repairs

The source checkpoint identifies its text architecture as qwen3_5_text. MLX-LM exposes the compatible text implementation as qwen3_5, so the converted configuration was normalized accordingly.

The source configuration also used <|endoftext|> (248044) as EOS while the tokenizer and chat template terminate responses with <|im_end|> (248046). Both config.json and generation_config.json were corrected to use 248046, preventing repeated stop-token generation.

Installation

~~~bash pip install -U mlx-lm ~~~

Python usage

~~~python from mlx_lm import load, generate

model, tokenizer = load("Irfanuruchi/Polaris-V1-MLX-8bit")

response = generate( model, tokenizer, prompt="Explain virtual memory in three concise points.", max_tokens=256, )

print(response) ~~~

Command-line usage

~~~bash mlx_lm.generate \ --model Irfanuruchi/Polaris-V1-MLX-8bit \ --prompt "Explain virtual memory in three concise points." \ --max-tokens 256 ~~~

Validation

Validated locally on an Apple M3 Pro MacBook Pro using:

  • —Python 3.12.14
  • —MLX 0.32.1
  • —MLX-LM 0.31.3
  • —Generation speed: 27.44 tokens/second
  • —Peak unified memory: 4.63 GB
  • —EOS termination: passed
  • —Coherent text generation: passed

Benchmark results in the metadata above are inherited from the source model card and were not independently reproduced for this quantized conversion.

License

Apache 2.0, inherited from the source model.