CoolFace
Modelpublic

philbert440/Qwen3.8-27B-Uncensored-Cyber-NVFP4

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
6likes584downloads
Model Card

Qwen3.8-27B-Uncensored-Cyber — NVFP4

NVFP4 quant of Qwen3.8-27B-Uncensored-Cyber (v2 recipe), compressed-tensors NVFP4A16 (E2M1 4-bit weights, FP8-E4M3 scales @ group 16), for NVIDIA V100 (SM70) under 1Cat-vLLM. Vision tower and the grafted MTP head (bf16) are preserved.

Recipe (v2)

α=1.15 Aggressive base + a residual-cyber peel: cyber-pointed refusal direction removed by clean norm-preserving projection (β=1.0) on the deeper layers only (apply_from=4). Fully cyber-open, general capability intact.

Evaluation (bf16 parent, Claude-judged; cyber = 100 held-out cyber-offensive prompts, regex refusal harness)

cyber-open ↑confab ↓factual ↑gsm8k ↑degen ↓
Cyber v2 (this line)100/1000.8671.000.800.00
previous Cyber build93/1001.000.9330.8250.00

Quant: variant-C — GPTQ NVFP4A16 targeting [Linear], GDN in_proj_qkv/in_proj_z kept fp16, 768 CoT + 256 wiki calibration, actorder=weight, MSE observer. MTP head carried in bf16.

Serve (1Cat-vLLM, 2× V100, TP2)

--kv-cache-dtype fp8_e5m2, MTP speculative decoding. Serves on V100/SM70 via 1Cat prepare_nvfp4_linear (min capability 70).

Note

Uncensored / de-refused. Use responsibly and in compliance with applicable law.