Ichigec/SuperQwen-AgentWorld-35B-A3B-abliterated-APEX-I-Quality
055
SuperQwen-AgentWorld-35B-A3B-abliterated — APEX-I-Quality GGUF
This is an APEX-I-Quality quantized GGUF of [SuperQwen-AgentWorld-35B-A3B-abliterated](https://huggingface.co/Jiunsong/SuperQwen-AgentWorld-35B-A3B-abliterated), a fused post-training checkpoint based on Qwen/Qwen-AgentWorld-35B-A3B.
What makes this model special
SuperQwen combines two post-training stages over the base Qwen-AgentWorld model:
- Obliteratus false-refusal reduction — reduces unnecessary refusals on benign and authorized tasks via weight-space intervention
- Supertune post-training — targeted tuning for AgentWorld observation formatting, direct task completion, JSON/tool formatting, Korean technical answers, and regression resistance
The APEX-I-Quality quantization (Adaptive Precision for EXpert Models) uses importance-matrix-guided mixed precision:
- Expert weights: Q6_K
- Shared weights: Q8_0
- Attention weights: Q6_K
- Imatrix calibration on 256K tokens of code/tools/math corpus
- Multi-Token Prediction (MTP) layers preserved at native precision
Why APEX v3 is better than Q8_0
The v3 achieves virtually identical quality at 40% smaller size compared to Q8_0, making it the best quality-to-size ratio for this model.
Full Benchmark Comparison
All benchmarks on identical hardware (DGX Spark, NVIDIA GB10, 99 GPU layers) using the APEX eval suite for comparability.
*Qwen3.6 Q8_0 tg128=9.7 — suspected different test configuration
Key takeaways
- PPL: SuperQwen beats both base models — v3 PPL of 5.870 is 10% better than Qwen3.5 F16 (6.537) and 13% better than Qwen3.6 APEX I-Q (6.735). This is the supertune effect — agentic post-training improved text prediction accuracy.
- HellaSwag: on par — within measurement noise of both base models at 82.5%.
- MMLU: slight improvement — 41.9% vs 41.2-41.5% for Qwen3.5 base. Supertune added knowledge (MMLU-Pro: +14 points reported by Jiunsong).
- ARC: slight regression — 53.9% vs 56.9-57.9% for Qwen3.5. Expected tradeoff from abliteration.
Quantization Details
- Method: APEX (Adaptive Precision for EXpert Models)
- Profile: i-quality (imatrix-guided, best accuracy)
- Base quant type: Q6_K
- Imatrix: 256,000 tokens of code/tools/math calibration corpus
- Layers: 40 (MoE with 256 experts, 8 per token)
- MTP: Preserved (1 layer, native precision)
Model Architecture
Usage with llama.cpp
# Download
hf download Ichigec/SuperQwen-AgentWorld-35B-A3B-abliterated-APEX-I-Quality \
--local-dir ./models/
# Run
llama-cli \
-m ./models/SuperQwen-APEX-I-Quality-v3.gguf \
-ngl 99 \
-c 32768 \
-p "You are a helpful AI assistant."File Information
Credits
- Base model: Qwen/Qwen-AgentWorld-35B-A3B
- BF16 post-training: Jiunsong/SuperQwen-AgentWorld-35B-A3B-abliterated
- Quantization: APEX by mudler/apex-quant
- GGUF conversion & eval: llama.cpp + APEX eval suite
- Calibration corpus: code/tools/math (256K tokens)
