CoolFace
Modelpublic

Ichigec/SuperQwen-AgentWorld-35B-A3B-abliterated-APEX-I-Quality

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes55downloads
Model Card

SuperQwen-AgentWorld-35B-A3B-abliterated — APEX-I-Quality GGUF

This is an APEX-I-Quality quantized GGUF of [SuperQwen-AgentWorld-35B-A3B-abliterated](https://huggingface.co/Jiunsong/SuperQwen-AgentWorld-35B-A3B-abliterated), a fused post-training checkpoint based on Qwen/Qwen-AgentWorld-35B-A3B.

What makes this model special

SuperQwen combines two post-training stages over the base Qwen-AgentWorld model:

  1. 1.Obliteratus false-refusal reduction — reduces unnecessary refusals on benign and authorized tasks via weight-space intervention
  2. 2.Supertune post-training — targeted tuning for AgentWorld observation formatting, direct task completion, JSON/tool formatting, Korean technical answers, and regression resistance

The APEX-I-Quality quantization (Adaptive Precision for EXpert Models) uses importance-matrix-guided mixed precision:

  • —Expert weights: Q6_K
  • —Shared weights: Q8_0
  • —Attention weights: Q6_K
  • —Imatrix calibration on 256K tokens of code/tools/math corpus
  • —Multi-Token Prediction (MTP) layers preserved at native precision

Why APEX v3 is better than Q8_0

MetricSuperQwen Q8_0 (35 GB)**APEX v3 (21 GB)**Δ
Perplexity5.8375.870+0.033 (within noise)
Size35 GB21 GB−40% smaller
Speed (tg128)—40.2 t/s—

The v3 achieves virtually identical quality at 40% smaller size compared to Q8_0, making it the best quality-to-size ratio for this model.

Full Benchmark Comparison

All benchmarks on identical hardware (DGX Spark, NVIDIA GB10, 99 GPU layers) using the APEX eval suite for comparability.

MetricQwen3.5 F16Qwen3.5 Q8_0Qwen3.5 APEX I-QQwen3.6 Q8_0Qwen3.6 APEX I-Q**Our Q8_0****Our APEX v1****Our APEX v3**
Size65 GB34 GB21 GB~34 GB21 GB35 GB22 GB21 GB
PPL6.5376.5336.5526.7206.7355.8375.8685.870
HellaSwag82.5%83.0%83.5%82.5%82.5%—82.75%82.50%
Winogrande74.5%75.3%74.5%———75.50%75.50%
MMLU41.5%41.2%41.4%———42.38%41.93%
ARC56.9%57.9%57.9%———54.52%53.85%
tg128 (t/s)30.452.563.19.7*63.9—38.540.2
*Qwen3.6 Q8_0 tg128=9.7 — suspected different test configuration

Key takeaways

  1. 1.PPL: SuperQwen beats both base models — v3 PPL of 5.870 is 10% better than Qwen3.5 F16 (6.537) and 13% better than Qwen3.6 APEX I-Q (6.735). This is the supertune effect — agentic post-training improved text prediction accuracy.
  2. 2.HellaSwag: on par — within measurement noise of both base models at 82.5%.
  3. 3.MMLU: slight improvement — 41.9% vs 41.2-41.5% for Qwen3.5 base. Supertune added knowledge (MMLU-Pro: +14 points reported by Jiunsong).
  4. 4.ARC: slight regression — 53.9% vs 56.9-57.9% for Qwen3.5. Expected tradeoff from abliteration.

Quantization Details

  • —Method: APEX (Adaptive Precision for EXpert Models)
  • —Profile: i-quality (imatrix-guided, best accuracy)
  • —Base quant type: Q6_K
  • —Imatrix: 256,000 tokens of code/tools/math calibration corpus
  • —Layers: 40 (MoE with 256 experts, 8 per token)
  • —MTP: Preserved (1 layer, native precision)

Model Architecture

ParameterValue
ArchitectureQwen3.5MoE
Total parameters35B
Activated parameters3B
Layers40 (10 full attention + 30 linear attention)
Experts256 (top-8 per token)
Context length262,144 tokens
Vocabulary248,320
MTP layers1

Usage with llama.cpp

bash
# Download
hf download Ichigec/SuperQwen-AgentWorld-35B-A3B-abliterated-APEX-I-Quality \
  --local-dir ./models/

# Run
llama-cli \
  -m ./models/SuperQwen-APEX-I-Quality-v3.gguf \
  -ngl 99 \
  -c 32768 \
  -p "You are a helpful AI assistant."

File Information

FileSizeSHA256
SuperQwen-APEX-I-Quality-v3.gguf21.3 GBafb6a47af5301e45d7e1de792e76d5611acb9d909b2a6d2a69fd645a12052162

Credits