CoolFace
Modelpublic

trevor000/Qwen3.6-35B-A3B-I-Mini-RTX4070

sourceHugging Faceapache-2.0updated 7d agoView on Hugging Face
0likes169downloads
Model Card

Qwen3.6-35B-A3B I-Mini, RTX 4070 research mirror

This is an attributed mirror of mudler/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-APEX-MTP-GGUF, using the I-Mini GGUF named in the manifest. The weight is unchanged upstream content; this project did not train it and makes no ownership claim over the base or APEX work.

Provenance

The model card identifies the base lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled, APEX quantization by the LocalAI/APEX team, and bundled Qwen MTP. The 14,272,221,568-byte I-Mini artifact has been verified against the SHA256 in provenance.json. The source revision is acb957d8dd1e806dadefc056daeb4ccf4d329287.

Download the reviewed file with hf download trevor000/Qwen3.6-35B-A3B-I-Mini-RTX4070 --local-dir Qwen3.6-35B-A3B-I-Mini-RTX4070. Use a llama.cpp build with the APEX/MTP support documented by the upstream model card, and see the RTX 4070 research repository for the recovered launcher context; this mirror does not bundle a runtime.

Recovered RTX 4070 baseline, 2026-09-20

On llama.cpp 3f545beccee69d9975f466ec7e45fd9aacd8ba90, three 512-token temperature-zero runs measured 94.656 tok/s median server decode and 88.386 tok/s median end-to-end wall throughput. This used unchanged top-k8 routing, Q8 KV, 16K allocated context, CPU placement for expert tensors in blocks 33–40, and batch/microbatch 1024/128. It is an initial recovered speed control, not a broad quality evaluation.

A corrected-fact retrieval check passed at 12,564 input tokens and again with actual conversation history at 12,766 tokens, decoding around 86–87 tok/s. Cold-history prefill took about 26 seconds. These are narrow checks; they do not establish general intelligence or perfect long-context memory. Further cache/prefetch experiments remain separate from this baseline. Historical reduced-top-k results are approximate routing changes and are not the recommended default.

Attribution and license

The source card declares Apache-2.0. Retain the Qwen Apache license and identify the lordx64 distillation, APEX/LocalAI quantization, Qwen MTP, and llama.cpp MTP contributions. This mirror is not an official Qwen, LocalAI, or APEX release.