CoolFace
Modelpublic

Abiray/Qwen3-30B-A3B-Instruct-2507-REAM-heretic-GGUF

sourceHugging Faceupdated 7mo agoView on Hugging Face
0likes546downloads
Model Card

Qwen3-30B-A3B-Instruct-2507-REAM-heretic - GGUF

This repository contains GGUF format model files for sasa2000/Qwen3-30B-A3B-Instruct-2507-REAM-heretic.

These files were quantized using llama.cpp to make them runnable on consumer hardware via CPU and GPU offloading.

Model Architecture Details

  • —Architecture: Qwen3 MoE (Mixture of Experts)
  • —Total Parameters: ~30B
  • —Active Parameters: ~3.3B (A3B)
  • —Max Context Length: 256K (Note: Requires significant RAM/VRAM at full context; lowering to 8K-32K is recommended for standard hardware).

Available Quantizations

File NameBit DepthDescription
[Qwen3-30B-A3B-Instruct-2507-REAM-heretic-Q4_K_M.gguf](https://huggingface.co/Abhiray/Qwen3-30B-A3B-Instruct-2507-REAM-heretic-GGUF/blob/main/Qwen3-30B-A3B-Instruct-2507-REAM-heretic-Q4_K_M.gguf)4-bit⭐ Recommended. Best balance of speed, RAM usage, and intelligence.
[Qwen3-30B-A3B-Instruct-2507-REAM-heretic-Q5_K_M.gguf](https://huggingface.co/Abhiray/Qwen3-30B-A3B-Instruct-2507-REAM-heretic-GGUF/blob/main/Qwen3-30B-A3B-Instruct-2507-REAM-heretic-Q5_K_M.gguf)5-bitHigher quality, slightly larger RAM requirement.
[Qwen3-30B-A3B-Instruct-2507-REAM-heretic-Q6_K.gguf](https://huggingface.co/Abhiray/Qwen3-30B-A3B-Instruct-2507-REAM-heretic-GGUF/blob/main/Qwen3-30B-A3B-Instruct-2507-REAM-heretic-Q6_K.gguf)6-bitVery high quality, near unquantized performance.
[Qwen3-30B-A3B-Instruct-2507-REAM-heretic-Q8_0.gguf](https://huggingface.co/Abhiray/Qwen3-30B-A3B-Instruct-2507-REAM-heretic-GGUF/blob/main/Qwen3-30B-A3B-Instruct-2507-REAM-heretic-Q8_0.gguf)8-bitExtremely high quality, large file size.
[Qwen3-30B-A3B-Instruct-2507-REAM-heretic-Q3_K_M.gguf](https://huggingface.co/Abhiray/Qwen3-30B-A3B-Instruct-2507-REAM-heretic-GGUF/blob/main/Qwen3-30B-A3B-Instruct-2507-REAM-heretic-Q3_K_M.gguf)3-bitSmallest file size, noticeable intelligence loss. Use only if heavily RAM constrained.

Prompt Format (ChatML)

This model uses the ChatML format. If you are using a UI like LM Studio or Ollama, it should detect this automatically from the GGUF metadata.

text
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
Hello!<|im_end|>
<|im_start|>assistant