Abiray/Qwen3-30B-A3B-Instruct-2507-REAM-heretic-GGUF
0546
Qwen3-30B-A3B-Instruct-2507-REAM-heretic - GGUF
This repository contains GGUF format model files for sasa2000/Qwen3-30B-A3B-Instruct-2507-REAM-heretic.
These files were quantized using llama.cpp to make them runnable on consumer hardware via CPU and GPU offloading.
Model Architecture Details
- Architecture: Qwen3 MoE (Mixture of Experts)
- Total Parameters: ~30B
- Active Parameters: ~3.3B (A3B)
- Max Context Length: 256K (Note: Requires significant RAM/VRAM at full context; lowering to 8K-32K is recommended for standard hardware).
Available Quantizations
Prompt Format (ChatML)
This model uses the ChatML format. If you are using a UI like LM Studio or Ollama, it should detect this automatically from the GGUF metadata.
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
Hello!<|im_end|>
<|im_start|>assistant