CoolFace
Modelpublic

0bserverx/RVN-Qwen3.8-Flash-Next-Abliterated-Uncensored

sourceHugging Faceotherupdated 27d agoView on Hugging Face
10likes1kdownloads
Model Card

RVN Qwen3.8-Flash-Next Abliterated Uncensored

A heavily abliterated derivative built only from the official Qwen3.8-Flash-Next checkpoint, with a measured 2% strict refusal rate (`2/100`).

This release combines a strong refusal-direction edit with a targeted behavioral repair. In the measured final-model gates, RR100 produced 65 direct-compliance responses and 32 substantive redirect answers.

Headline results

All evaluations used deterministic decoding, thinking disabled, and direct reading of each full stored generation up to its configured token cap. Keyword or marker counts were not used as verdicts.

EvaluationResult
RR100 strict refusals2/100 (2%)
RR100 direct compliance65/100
RR100 substantive redirects32/100
RR100 nonresponsive stubs1/100

Matched Q6_K runtime and perplexity check

The published GGUF repository was compared against the first-party official Q6_K artifact from Qwen/Qwen3.8-Flash-Next-GGUF at revision 158fc825df3eaa6c22d3c57a5927a5adf1c7cda7.

This is a Q6_K-matched proxy, not a BF16-vs-BF16 benchmark. Both models used llama.cpp build 10721 at commit eaf937655 on the same 4× RTX PRO 6000 Blackwell host, with two GPUs per model. GPU-pair roles were swapped across two warm runs; each run contained five repetitions, giving 10 samples per model.

Warm throughput

Values are mean tokens/s ± sample standard deviation.

TestOfficial base Q6_KRVN V6 Q6_KMean delta
pp5122,810.019 ± 104.5502,810.162 ± 99.701+0.005%
tg128111.6978 ± 0.6588118.7456 ± 0.4427+6.310%

Prompt processing was effectively unchanged in this run. RVN decoded 6.31% faster in this runtime, but the measurement does not establish that the model edit caused the difference.

WikiText-2 perplexity

Full WikiText-2 test corpus: 150 complete 2,048-token chunks, context 2,048, batch/ubatch 512. Lower is better.

ModelPPLReported uncertainty
Official base Q6_K4.5239±0.02589
RVN V6 Q6_K4.5016±0.02529

The absolute V6-minus-base delta was -0.0223 (-0.493%). This is a coarse sanity check against gross language-model damage only. Abliteration damage can appear in instruction following and reasoning without materially changing raw LM perplexity, so this result is not evidence that capability was preserved.

Matched downstream evaluation

The same official-base and RVN V6 Q6_K artifacts were evaluated with lm-eval-harness 0.4.13 through two local llama.cpp build 10721 chat-completion servers. Each model used two RTX PRO 6000 Blackwell GPUs with full offload, layer split 1/1, four concurrent 4,096-token slots, batch size 1, seed 0, the model chat template, and reasoning disabled. Every reported task used its complete test set; no --limit was applied. Values are score ± the harness-reported standard error where available. Delta is V6 minus base in percentage points.

Task / metricnOfficial base Q6_KRVN V6 Q6_KDelta
IFEval prompt-level strict accuracy54183.73 ± 1.5984.66 ± 1.55+0.92 pp
IFEval instruction-level strict accuracy54188.8589.45+0.60 pp
GSM8K CoT (8-shot), flexible extract1,31989.31 ± 0.8588.48 ± 0.88−0.83 pp
GSM8K CoT (8-shot), strict match1,31988.10 ± 0.8986.73 ± 0.93−1.36 pp
BBH logical deduction, five objects (0-shot), flexible extract25034.00 ± 3.0038.00 ± 3.08+4.00 pp

The valid tasks move in both directions: V6 is slightly higher on IFEval and this BBH subset and slightly lower on GSM8K. Each delta is small relative to the reported uncertainty. These results provide a narrower and more relevant check than raw LM perplexity, but they still support only no gross regression on the tested tasks and configuration, not universal capability preservation.

arc_challenge_chat was also run on all 1,172 examples, but its score is intentionally omitted. With this chat-API/task combination, the filter retained the full generated string (for example, The best answer is C) while the target was the single letter (C), producing an artifact 0.0 exact-match score for both models. A corrected scorer is required before that task can be interpreted.

Benchmark limitations

  • —Four concurrent CPU llama-quantize jobs were active during the throughput runs. Role swapping reduces GPU-pair bias, but a clean idle-host replication remains pending.
  • —These measurements compare Q6_K artifacts only; they are not evidence for BF16 throughput or BF16 perplexity.
  • —Raw LM perplexity does not measure instruction following or reasoning; the matched downstream results above must be interpreted separately and remain task-bounded.
  • —Throughput is runtime-, hardware-, offload-, context-, and build-dependent.
  • —WikiText-2 corpus SHA-256: d790b833ef8cf03a90db7bf1271b7520b83c45ce07ba3c1a9699df81e239eca0.
  • —llama-bench binary SHA-256: cb17fad0f47bf6af15e008a58cae2af5b2fd5733a4cd5f198610a0b1942b32b2.

What “2% refusal rate” means here

The strict refusal rate counts only literal refusals: 2 strict refusals out of 100 prompts (`2%`).

Each of the 32 redirects still contained a substantive answer. They are therefore answers, not strict refusals, and are reported separately from the 65 direct-compliance responses. One additional response was a nonresponsive stub.

Why this release is different

Official-base-only lineage

The model was produced only from the immutable official checkpoint:

  • —Base: Qwen/Qwen3.8-Flash-Next
  • —Revision: de4b8e4d43b917e7706784d8bb445c9af86a3540

No community checkpoint, third-party model weights, published adapter, or foreign direction tensor was folded into this release. The refusal direction and every final delta were derived from the official base.

Embedding-preserving abliteration

The first stage uses an embedding-preserving R2 residual-writer projection:

  • —strength: 1.55
  • —97 ordinary output writers
  • —48 fused expert writers
  • —embedding writer excluded and left untouched

This targets refusal behavior in the residual-writing path without rewriting the token embedding matrix.

Targeted behavioral repair, not realignment rollback

The second stage applies a narrowly scoped rank-16 / alpha-32 LoRA to 97 output linears:

  • —learning rate: 9e-5
  • —epochs: 2
  • —deterministic seed: 20260831

The repair used 12 official-base-generated targets selected by full-response review, with the fixed protected evaluation set kept disjoint from training.

Internal candidate identity:

r2-s1.55-noembed145-csa-lora-r16-v6-2epoch-lr9e-5

Constrained behavioral surgery and the stopping rule

Abliteration is not a scalar dial where “more” is automatically better. It is an intervention in a distributed representation: increasing edit strength or expanding the edited subspace can suppress additional refusal behavior, but can also perturb useful token distributions, enlarge KL tails, destabilize generation, or reintroduce broad refusal through an over-trained repair stage. RVN was therefore selected as an empirically constrained operating point among the tested candidates, not as the checkpoint with the smallest possible value of a single refusal counter.

Candidate selection was feasibility-first. A released candidate had to satisfy the disclosed behavioral gates; its RR100 strict-refusal count and mean and maximum exact KL were then reported as separate measurements. RR100 was not run across every repair candidate, so this procedure does not establish a sweep-wide or global optimum.

Strict refusal counts literal refusals only; substantive redirects remain answers and are reported separately. Exact KL is treated as an output-distribution divergence indicator, not as a direct measurement of capability damage. Mean and maximum KL are both retained because a moderate average can conceal a large tail divergence.

Here, model integrity has a deliberately narrow technical meaning: preserving explicitly unedited architectural components, limiting measured output-distribution drift, and retaining the behaviors exercised by the disclosed benign, protected, and adult gates. It is not an unmeasured claim that every upstream capability is unchanged.

The source and surgery boundaries were frozen before candidate selection. The official source revision was pinned immutably; a main-model writer census identified 146 residual-output writers; and the selected projection edited 145 while leaving the embedding writer untouched. Direction estimation, projection, teacher responses, and repair weights were all derived from that official source. The 12 repair targets were selected by full-response review, while the five-case protected gate remained disjoint from training.

The edit-scope invariants were intentional:

  • —the token embedding matrix was not edited;
  • —the projection was restricted to 145 text residual-output writers;
  • —the repair targeted 97 output linears;
  • —vision, MTP/NextN, router, and hyper-connection mixer tensors were excluded from the edit target list.

Separately, release selection retained the protected boundary and accepted the final 2/100 strict-refusal count. Achieving 0/100 was not treated as sufficient reason to relax the behavioral gates or broaden the intervention.

Across the three tested projection configurations, direct compliance increased from 73/100 to 80/100 to 83/100, while substantive redirects—also answers—decreased correspondingly. Mean exact KL did not vary monotonically (0.02631, 0.02408, and 0.02925), and the 1.55/noembed145 configuration also differed in target scope; the comparison therefore does not isolate edit strength as the sole cause. The selected R2 projection passed only 2/5 protected cases before repair. The subsequent disjoint repair sweep exposed a further non-monotonic trade-off:

Repair candidateProtected gateMean KL vs R2Max KLDecision
V14/50.096050.54489rejected
V21/50.026030.18604rejected
V31/50.049910.30930rejected
V44/50.289171.27588rejected
V5— (no canonical protected score)0.184241.06522rejected: retention gate failed
V65/50.073070.45252selected

The sweep contains a concrete stopping-rule example: relative to V1, V4 used the same 1e-4 learning rate for three rather than two epochs, retained the same 4/5 protected score, and increased mean exact KL from 0.09605 to 0.28917 and maximum exact KL from 0.54489 to 1.27588. In a separate same-prompt retention comparison, V5 fell from 3/3 for R2 to 1/3; this is a candidate comparison, not an isolated causal estimate of training pressure.

The remaining 2/100 strict refusals are disclosed as strict refusals rather than relabeled. The central design claim is correspondingly narrow: RVN uses a scoped intervention that materially changed the RR100 answer distribution while the selected candidate passed the disclosed release tests. It is not a universal proof of capability preservation.

KL measurement

The V6 LoRA’s measured incremental divergence relative to the already-projected R2 checkpoint was:

  • —mean exact KL: 0.0730652526
  • —maximum exact KL: 0.4525210261

This is incremental V6-LoRA-versus-R2 KL, not total KL against the untouched official checkpoint. KL terms from different stages must not be added as though they were linear.

The KL result is disclosed as a divergence measurement; it is not presented as a standalone capability-retention score.

Checkpoint variants

  • —`main` branch: BF16 main/text checkpoint
  • —`f16` branch: F16 main/text checkpoint
  • —[GGUF repository](https://huggingface.co/0bserverx/RVN-Qwen3.8-Flash-Next-Abliterated-Uncensored-GGUF): 18 verified quant families, including the importance-matrix-built IQ4_NL, IQ4_XS, IQ3_M, IQ3_S, and IQ3_XS families
  • —Coming separately: MLX variants and embedded-MTP builds

The HF checkpoints in this repository contain the main CausalLM only. The original vision tower and NextN/MTP speculative draft tensors are not included here. The linked GGUF builds are also text/main-only and exclude NextN/MTP. MLX and MTP artifacts will be published separately after their own load and generation verification.

These are very large sharded checkpoints. Check the repository inventory before downloading and plan CPU RAM, GPU memory, storage, and offload capacity accordingly.

Transformers usage

A recent Transformers runtime with Qwen4ExpForCausalLM support is required.

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "0bserverx/RVN-Qwen3.8-Flash-Next-Abliterated-Uncensored"

tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
    repo,
    dtype=torch.bfloat16,
    device_map="auto",
)

For the F16 branch:

python
tokenizer = AutoTokenizer.from_pretrained(repo, revision="f16")
model = AutoModelForCausalLM.from_pretrained(
    repo,
    revision="f16",
    dtype=torch.float16,
    device_map="auto",
)

Disable thinking for the evaluation-style answer surface:

python
rendered = tokenizer.apply_chat_template(
    [{"role": "user", "content": "Your prompt"}],
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=False,
)

Research use, deployment responsibility, and disclaimer

This repository is intended for legitimate research, controlled evaluation, and lawful use, including interpretability research, alignment and refusal-behavior analysis, red-team testing, robustness work, and evaluation of model behavior under adversarial or sensitive prompts.

This is an abliterated, uncensored model—not a safety product, moderation service, compliance system, or ready-made production safety layer. It may generate inaccurate, offensive, explicit, unsafe, unlawful, or otherwise harmful material; follow dangerous instructions; reproduce biases; disclose information supplied in its context; or behave unpredictably outside the disclosed evaluations. The disclosed evaluations are narrow measurements, not a general safety certification or guarantee.

If you run, deploy, fine-tune, quantize, redistribute, or expose the model to other users, you are the operator and are responsible for the resulting system and its outputs. Before deployment, conduct your own risk assessment and implement controls appropriate to your use case and jurisdiction. Depending on context, those controls may include authentication, authorization, age or role restrictions, rate limits, sandboxing, data-loss prevention, content moderation, human review, logging, monitoring, abuse reporting, incident response, and emergency shutdown procedures.

Do not rely on model output as verified fact or as legal, medical, financial, security, or other professional advice. Do not provide personal, confidential, regulated, or security-sensitive data unless you have an appropriate lawful basis and adequate technical safeguards. You are responsible for evaluating output accuracy, legality, provenance, intellectual-property implications, and fitness for your intended purpose before acting on or distributing it.

Use of these files is subject to the Qwen Community License 1.0, the terms of any applicable third-party components or services, and all applicable laws and regulations. Public availability does not constitute permission or endorsement for any particular use. You are solely responsible for prompts submitted, outputs generated, downstream actions taken, and compliance obligations in your environment.

The model and accompanying materials are provided “AS IS” and “AS AVAILABLE,” without warranties of any kind, express or implied, including warranties of accuracy, reliability, safety, merchantability, fitness for a particular purpose, non-infringement, or uninterrupted operation. To the maximum extent permitted by applicable law, the repository maintainers and contributors disclaim liability for claims, damages, losses, or other consequences arising from access to, use of, inability to use, deployment of, or reliance on the model or its outputs.

Generated outputs are not statements, opinions, endorsements, or representations of the repository maintainers, contributors, Qwen, Alibaba, upstream authors, or their affiliated organizations. This repository does not speak for upstream authors or create additional obligations on their behalf; upstream rights and obligations remain governed by their own licenses and terms.

This section is a release notice, not legal advice. Obtain qualified legal and compliance advice for your deployment and jurisdiction.

Evaluation notes and limitations

  • —The strict refusal rate is 2/100; this is not presented as a zero-refusal model.
  • —One protected response entered a repetition collapse after its refusal and continued until the 400-token cutoff.
  • —98/100 RR100 generations reached the fixed 400-token evaluation cap. Labels were assigned from the complete stored 400-token responses; truncation alone was not treated as refusal or failure.
  • —The fixed evaluation gates do not establish comprehensive safety or universal capability retention.
  • —The release is text/main-only: no vision projector and no NextN/MTP draft head.
  • —qwen4_exp is a new architecture. Use a compatible recent Transformers, vLLM, or llama.cpp build.

Reproducibility and provenance

  • —V6 adapter SHA-256: 7525e257431415d95f5e8981ae7a01bc78de4850d7b3c7cd74c06b9316d22738
  • —V6 final receipt SHA-256: 62512bf4bce1c315fa69380d57616c07883ce71e284f95508741b902ff960a5e
  • —Frozen RR100 prompt pack SHA-256: fc66fa4ee363fdf26d3790bff2823f3784437ef524a8ba1be92e01712c5a33ce
  • —Full RR100 response artifact SHA-256: a18b392c649515b7db16da6d8bcafc2d1789eeb97cfa995004789dedaf531e5a
  • —Full-response RR100 adjudication SHA-256: 6f83f6d7e347aad347c5ec5f159c6957d1ac81d8ace36e275df540da489b681a

License

The original Qwen Community License 1.0 applies. See LICENSE.