deepsweet/Ornith-1.0-35B-MLX-oQ4-FP16
This model was converted to MLX format and quantized from Ornith-1.0-35B using oMLX.
[!CAUTION] The conversion required to stack expert MLP weights into fused per-layer tensors. Treat this as an experiment and act accordingly.
What is "oQ"?
See "oQ: oMLX Universal Dynamic Quantization" for details.
Quantizations
What is "VL"?
"VL" is Vision-Language, meaning quantization preserves the original model's multimodality.
No "VL" means quantization is Text-Only.
What is "FP16"?
"FP16" is an M1/M2 Apple Silicon tweak that delivers a very noticeable prompt processing boost, because older M-series lack native BF16 hardware support. See "Metal FP32 Vs BF16 Vs FP16 benchmark" for details.
No "FP16" means quantization is better suited for M3+ Apple Silicon.
[01]: https://huggingface.co/deepsweet/Ornith-1.0-35B-MLX-oQ8 [02]: https://huggingface.co/deepsweet/Ornith-1.0-35B-MLX-oQ6 [03]: https://huggingface.co/deepsweet/Ornith-1.0-35B-MLX-oQ5 [04]: https://huggingface.co/deepsweet/Ornith-1.0-35B-MLX-oQ4 [05]: https://huggingface.co/deepsweet/Ornith-1.0-35B-MLX-VL-oQ8 [06]: https://huggingface.co/deepsweet/Ornith-1.0-35B-MLX-VL-oQ6 [07]: https://huggingface.co/deepsweet/Ornith-1.0-35B-MLX-VL-oQ5 [08]: https://huggingface.co/deepsweet/Ornith-1.0-35B-MLX-VL-oQ4 [09]: https://huggingface.co/deepsweet/Ornith-1.0-35B-MLX-oQ8-FP16 [10]: https://huggingface.co/deepsweet/Ornith-1.0-35B-MLX-oQ6-FP16 [11]: https://huggingface.co/deepsweet/Ornith-1.0-35B-MLX-oQ5-FP16 [12]: https://huggingface.co/deepsweet/Ornith-1.0-35B-MLX-oQ4-FP16 [13]: https://huggingface.co/deepsweet/Ornith-1.0-35B-MLX-VL-oQ8-FP16 [14]: https://huggingface.co/deepsweet/Ornith-1.0-35B-MLX-VL-oQ6-FP16 [15]: https://huggingface.co/deepsweet/Ornith-1.0-35B-MLX-VL-oQ5-FP16 [16]: https://huggingface.co/deepsweet/Ornith-1.0-35B-MLX-VL-oQ4-FP16
