LibraxisAI/QwQ-32B-MLX-Q5
QwQ-32B-MLX-Q5
QwQ-32B-MLX-Q5 is an MLX Q5 checkpoint derived from Qwen/QwQ-32B, intended for local text generation on Apple Silicon.
Intended use
- Local text generation and chat-style prompting on Apple Silicon
- MLX-LM experimentation with the declared upstream model family
- Offline or operator-controlled inference workflows
Out of scope
- Safety-critical decisions without domain expert review
- Claims of benchmark superiority not backed by published evaluation data
- Non-MLX runtime guarantees; this card documents the shipped HF checkpoint, not every possible serving stack
Training and conversion metadata
This card only reports metadata present in the Hugging Face repository, existing card frontmatter, or public config files. Missing benchmark, dataset, or training-run details are left explicit rather than reconstructed.
Tested inference path
Inference for this checkpoint has been tested with [`LibraxisAI/mlx-batch-server`](https://github.com/LibraxisAI/mlx-batch-server).\ This is the recommended tested path for operator-controlled local inference on Apple Silicon.
This does not claim compatibility with every possible serving stack. It documents the path that has been exercised for this published checkpoint.
Usage
CLI
pip install mlx-lm
mlx_lm.generate \
--model LibraxisAI/QwQ-32B-MLX-Q5 \
--prompt "Summarize the key signals in this document and list the next action items." \
--max-tokens 400Python
from mlx_lm import load, generate
model, tokenizer = load("LibraxisAI/QwQ-32B-MLX-Q5")
prompt = "Summarize the key signals in this document and list the next action items."
response = generate(model, tokenizer, prompt=prompt, max_tokens=400)
print(response)Multi-turn with the chat template
This checkpoint follows the tokenizer/chat-template contract inherited from Qwen/QwQ-32B when the template is present in the repository:
from mlx_lm import load, generate
model, tokenizer = load("LibraxisAI/QwQ-32B-MLX-Q5")
messages = [
{"role": "user", "content": "Summarize the key signals in this document and list the next action items."},
]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
response = generate(model, tokenizer, prompt=prompt, max_tokens=400)
print(response)Example output
No public sample output is currently declared for this checkpoint.
Quantization notes
Limitations
- No public benchmarks for this checkpoint are declared in the model metadata.
- No public benchmark claims are made by this card unless listed in the frontmatter.
- Validate outputs on your own domain data before relying on this checkpoint.
- Memory use and speed depend heavily on the exact Apple Silicon generation, unified-memory size, and prompt length.
License
apache-2.0. Check the upstream/base model license as well when a base model is declared.
Citation
@misc{libraxisai-qwq-32b-mlx-q5,
title = {QwQ-32B-MLX-Q5},
author = {LibraxisAI},
year = {2026},
howpublished = {\url{https://huggingface.co/LibraxisAI/QwQ-32B-MLX-Q5}},
note = {MLX checkpoint published by LibraxisAI}
}𝚅𝚒𝚋𝚎𝚌𝚛𝚊𝚏𝚝𝚎𝚍. with AI Agents by VetCoders (c)2024-2026 LibraxisAI
