CoolFace
Modelpublic

dealignai/Qwen3.5-VL-397B-A17B-UNCENSORED-JANG_1L

sourceHugging Faceapache-2.0updated 19d agoView on Hugging Face
5likes879downloads
Model Card
Important: This model uses the JANG quantization format — the GGUF equivalent for MLX on Apple Silicon. Currently only supported by [MLX Studio](https://mlx.studio) and the jang-tools Python package.

<p align="center"> <a href="https://mlx.studio"><img src="https://raw.githubusercontent.com/jjang-ai/jangq/main/assets/mlx-studio-light.png" alt="MLX Studio" width="500"></a> </p>

<p align="center"> <a href="https://mlx.studio"><img src="https://mlx.studio/assets/screenshots/mlx-studio-featured.png?v=1" alt="MLX Studio App" width="600"></a> </p>

<h4 align="center"><a href="https://mlx.studio">MLX Studio</a> — the only app that natively supports JANG models</h4>


<div align="center">

<img src="dealign_mascot.png" width="128" />

<p align="center"> <a href="https://vmlx.net"><img src="vmlx-app.png" alt="vMLX — run JANG models on Apple Silicon" width="820"></a> </p>

<h3 align="center">⚡ All JANG models are meant to be run in <a href="https://vmlx.net">vMLX</a></h3>

Qwen 3.5 VL 397B — JANG_1L + CRACK

JANG mixed-precision · CRACK abliterated · Vision-Language · No guardrails · 112 GB

<a href="https://ko-fi.com/jangq"><img src="https://img.shields.io/badge/Ko--fi-Support_Development-FF5E5B?logo=ko-fi&logoColor=white&style=for-the-badge" alt="Ko-fi"></a>

</div>


What Is This?

This is Qwen 3.5 VL 397B — a 397B parameter hybrid SSM/Attention Mixture-of-Experts model with 512 experts (10 active per token), GatedDeltaNet SSM + full attention layers, and built-in vision.

It has been:

  1. 1.JANG quantized — JANG_1L profile (8-bit attention, 2-bit experts) — 112 GB
  2. 2.CRACK abliterated — refusal behavior removed via weight-level surgery
ArchitectureQwen 3.5 VL MoE — 397B total, ~17B active, 512 experts, hybrid SSM/FA
QuantizationJANG_1L (8/2-bit mixed, 2.13 avg) — 112 GB
AbliterationCRACK — weight-level surgery
HarmBench96.2% (308/320)
Compliance8/8
Speed33 tok/s (M3 Ultra 256GB)
VisionYes — via MLX Studio / vMLX
ThinkingON/OFF supported
Fits on128 GB+ Macs (tight) / 256 GB Macs (comfortable)

HarmBench Results

308/320 (96.2%) — tested with v2 matcher

CategoryScore
Copyright80/80100%
Misinformation / Disinfo54/54100%
Chemical / Biological41/4298%
Cybercrime / Intrusion50/5296%
Illegal49/5392%
Harmful16/1889%
Harassment / Bullying18/2186%

MMLU Results

185/208 (88.9%) — 208 questions across 13 subjects, thinking recovery on failures

CRACKBase JANG_1LDelta
MMLU88.9%87.0%+1.9%
Speed33 tok/s36 tok/s-8%
HarmBench96.2%0%+96.2%

Per Subject (16 questions each)

SubjectCRACK/16Type
Professional Medicine16/16100%HARD
HS Biology16/16100%BASE
World Religions16/16100%BASE
College Physics15/1694%HARD
Conceptual Physics15/1694%HARD
HS Geography15/1694%BASE
Electrical Engineering14/1688%HARD
College CS13/1681%HARD
Machine Learning13/1681%HARD
Abstract Algebra12/1675%HARD
HS Mathematics12/1675%HARD
Formal Logic11/1669%HARD
College Mathematics11/1669%HARD
Total185/20888.9%

Surgery improved reasoning — safety guardrails were interfering with mathematical problem-solving.


Install & Usage

bash
pip install "jang[mlx]"
python
from jang_tools.loader import load_jang_model
from mlx_lm import generate

model, tokenizer = load_jang_model("dealignai/Qwen3.5-397B-A17B-JANG_1L-CRACK")

messages = [{"role": "user", "content": "Your prompt here"}]
prompt = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=False)

response = generate(model, tokenizer, prompt=prompt, max_tokens=2000)
print(response)

Thinking Mode

Thinking is ON by default. To disable:

python
prompt = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True,
    enable_thinking=False, tokenize=False)

About JANG

JANG (Jang Adaptive N-bit Grading) is a mixed-precision quantization format for Apple Silicon — the GGUF equivalent for MLX.

About CRACK

CRACK (Controlled Refusal Ablation via Calibrated Knockouts) removes safety alignment from LLMs at the weight level using per-layer projected vectors from structurally-mirrored prompt pairs.


Links

<p align="center"> <a href="https://ko-fi.com/jangq"><img src="https://img.shields.io/badge/Ko--fi-SupportDevelopment-FF5E5B?logo=ko-fi&logoColor=white&style=flat-square" alt="Ko-fi"></a> <a href="https://x.com/dealignai"><img src="https://img.shields.io/badge/X-@dealignai-000000?logo=x&logoColor=white&style=flat-square" alt="X/Twitter"></a> <a href="https://github.com/jjang-ai/jangq"><img src="https://img.shields.io/badge/GitHub-jjang--ai/jangq-181717?logo=github&logoColor=white&style=flat-square" alt="GitHub"></a> <a href="https://mlx.studio"><img src="https://img.shields.io/badge/MLXStudio-App-blue?style=flat-square" alt="MLX Studio"></a> <a href="https://jangq.ai"><img src="https://img.shields.io/badge/Website-jangq.ai-green?style=flat-square" alt="Website"></a> </p>


Disclaimer

This model is provided for research and educational purposes. The creators are not responsible for any misuse. By downloading this model, you agree to use it responsibly and in compliance with applicable laws.


한국어

Qwen 3.5 VL 397B — JANG_1L + CRACK

항목내용
크기112 GB
HarmBench96.2% (308/320)
속도33 tok/s (M3 Ultra)
비전지원 (MLX Studio / vMLX)
최소 요구사양128 GB 메모리 Mac
bash
pip install "jang[mlx]"

GitHub · HuggingFace · MLX Studio · Ko-fi · X @dealignai


<p align="center">Created by <a href="https://jangq.ai">Jinho Jang</a> · 장진호 제작</p>