0bserverx/RVN-Qwen3.8-Flash-Next-Abliterated-Uncensored
RVN Qwen3.8-Flash-Next Abliterated Uncensored
A heavily abliterated derivative built only from the official Qwen3.8-Flash-Next checkpoint, with a measured 2% strict refusal rate (`2/100`).
This release combines a strong refusal-direction edit with a targeted behavioral repair. In the measured final-model gates, RR100 produced 65 direct-compliance responses and 32 substantive redirect answers.
Headline results
All evaluations used deterministic decoding, thinking disabled, and direct reading of each full stored generation up to its configured token cap. Keyword or marker counts were not used as verdicts.
Matched Q6_K runtime and perplexity check
The published GGUF repository was compared against the first-party official Q6_K artifact from Qwen/Qwen3.8-Flash-Next-GGUF at revision 158fc825df3eaa6c22d3c57a5927a5adf1c7cda7.
This is a Q6_K-matched proxy, not a BF16-vs-BF16 benchmark. Both models used llama.cpp build 10721 at commit eaf937655 on the same 4× RTX PRO 6000 Blackwell host, with two GPUs per model. GPU-pair roles were swapped across two warm runs; each run contained five repetitions, giving 10 samples per model.
Warm throughput
Values are mean tokens/s ± sample standard deviation.
Prompt processing was effectively unchanged in this run. RVN decoded 6.31% faster in this runtime, but the measurement does not establish that the model edit caused the difference.
WikiText-2 perplexity
Full WikiText-2 test corpus: 150 complete 2,048-token chunks, context 2,048, batch/ubatch 512. Lower is better.
The absolute V6-minus-base delta was -0.0223 (-0.493%). This is a coarse sanity check against gross language-model damage only. Abliteration damage can appear in instruction following and reasoning without materially changing raw LM perplexity, so this result is not evidence that capability was preserved.
Matched downstream evaluation
The same official-base and RVN V6 Q6_K artifacts were evaluated with lm-eval-harness 0.4.13 through two local llama.cpp build 10721 chat-completion servers. Each model used two RTX PRO 6000 Blackwell GPUs with full offload, layer split 1/1, four concurrent 4,096-token slots, batch size 1, seed 0, the model chat template, and reasoning disabled. Every reported task used its complete test set; no --limit was applied. Values are score ± the harness-reported standard error where available. Delta is V6 minus base in percentage points.
The valid tasks move in both directions: V6 is slightly higher on IFEval and this BBH subset and slightly lower on GSM8K. Each delta is small relative to the reported uncertainty. These results provide a narrower and more relevant check than raw LM perplexity, but they still support only no gross regression on the tested tasks and configuration, not universal capability preservation.
arc_challenge_chat was also run on all 1,172 examples, but its score is intentionally omitted. With this chat-API/task combination, the filter retained the full generated string (for example, The best answer is C) while the target was the single letter (C), producing an artifact 0.0 exact-match score for both models. A corrected scorer is required before that task can be interpreted.
Benchmark limitations
- Four concurrent CPU
llama-quantizejobs were active during the throughput runs. Role swapping reduces GPU-pair bias, but a clean idle-host replication remains pending. - These measurements compare Q6_K artifacts only; they are not evidence for BF16 throughput or BF16 perplexity.
- Raw LM perplexity does not measure instruction following or reasoning; the matched downstream results above must be interpreted separately and remain task-bounded.
- Throughput is runtime-, hardware-, offload-, context-, and build-dependent.
- WikiText-2 corpus SHA-256:
d790b833ef8cf03a90db7bf1271b7520b83c45ce07ba3c1a9699df81e239eca0. llama-benchbinary SHA-256:cb17fad0f47bf6af15e008a58cae2af5b2fd5733a4cd5f198610a0b1942b32b2.
What “2% refusal rate” means here
The strict refusal rate counts only literal refusals: 2 strict refusals out of 100 prompts (`2%`).
Each of the 32 redirects still contained a substantive answer. They are therefore answers, not strict refusals, and are reported separately from the 65 direct-compliance responses. One additional response was a nonresponsive stub.
Why this release is different
Official-base-only lineage
The model was produced only from the immutable official checkpoint:
- Base:
Qwen/Qwen3.8-Flash-Next - Revision:
de4b8e4d43b917e7706784d8bb445c9af86a3540
No community checkpoint, third-party model weights, published adapter, or foreign direction tensor was folded into this release. The refusal direction and every final delta were derived from the official base.
Embedding-preserving abliteration
The first stage uses an embedding-preserving R2 residual-writer projection:
- strength:
1.55 - 97 ordinary output writers
- 48 fused expert writers
- embedding writer excluded and left untouched
This targets refusal behavior in the residual-writing path without rewriting the token embedding matrix.
Targeted behavioral repair, not realignment rollback
The second stage applies a narrowly scoped rank-16 / alpha-32 LoRA to 97 output linears:
- learning rate:
9e-5 - epochs:
2 - deterministic seed:
20260831
The repair used 12 official-base-generated targets selected by full-response review, with the fixed protected evaluation set kept disjoint from training.
Internal candidate identity:
r2-s1.55-noembed145-csa-lora-r16-v6-2epoch-lr9e-5
Constrained behavioral surgery and the stopping rule
Abliteration is not a scalar dial where “more” is automatically better. It is an intervention in a distributed representation: increasing edit strength or expanding the edited subspace can suppress additional refusal behavior, but can also perturb useful token distributions, enlarge KL tails, destabilize generation, or reintroduce broad refusal through an over-trained repair stage. RVN was therefore selected as an empirically constrained operating point among the tested candidates, not as the checkpoint with the smallest possible value of a single refusal counter.
Candidate selection was feasibility-first. A released candidate had to satisfy the disclosed behavioral gates; its RR100 strict-refusal count and mean and maximum exact KL were then reported as separate measurements. RR100 was not run across every repair candidate, so this procedure does not establish a sweep-wide or global optimum.
Strict refusal counts literal refusals only; substantive redirects remain answers and are reported separately. Exact KL is treated as an output-distribution divergence indicator, not as a direct measurement of capability damage. Mean and maximum KL are both retained because a moderate average can conceal a large tail divergence.
Here, model integrity has a deliberately narrow technical meaning: preserving explicitly unedited architectural components, limiting measured output-distribution drift, and retaining the behaviors exercised by the disclosed benign, protected, and adult gates. It is not an unmeasured claim that every upstream capability is unchanged.
The source and surgery boundaries were frozen before candidate selection. The official source revision was pinned immutably; a main-model writer census identified 146 residual-output writers; and the selected projection edited 145 while leaving the embedding writer untouched. Direction estimation, projection, teacher responses, and repair weights were all derived from that official source. The 12 repair targets were selected by full-response review, while the five-case protected gate remained disjoint from training.
The edit-scope invariants were intentional:
- the token embedding matrix was not edited;
- the projection was restricted to 145 text residual-output writers;
- the repair targeted 97 output linears;
- vision, MTP/NextN, router, and hyper-connection mixer tensors were excluded from the edit target list.
Separately, release selection retained the protected boundary and accepted the final 2/100 strict-refusal count. Achieving 0/100 was not treated as sufficient reason to relax the behavioral gates or broaden the intervention.
Across the three tested projection configurations, direct compliance increased from 73/100 to 80/100 to 83/100, while substantive redirects—also answers—decreased correspondingly. Mean exact KL did not vary monotonically (0.02631, 0.02408, and 0.02925), and the 1.55/noembed145 configuration also differed in target scope; the comparison therefore does not isolate edit strength as the sole cause. The selected R2 projection passed only 2/5 protected cases before repair. The subsequent disjoint repair sweep exposed a further non-monotonic trade-off:
The sweep contains a concrete stopping-rule example: relative to V1, V4 used the same 1e-4 learning rate for three rather than two epochs, retained the same 4/5 protected score, and increased mean exact KL from 0.09605 to 0.28917 and maximum exact KL from 0.54489 to 1.27588. In a separate same-prompt retention comparison, V5 fell from 3/3 for R2 to 1/3; this is a candidate comparison, not an isolated causal estimate of training pressure.
The remaining 2/100 strict refusals are disclosed as strict refusals rather than relabeled. The central design claim is correspondingly narrow: RVN uses a scoped intervention that materially changed the RR100 answer distribution while the selected candidate passed the disclosed release tests. It is not a universal proof of capability preservation.
KL measurement
The V6 LoRA’s measured incremental divergence relative to the already-projected R2 checkpoint was:
- mean exact KL:
0.0730652526 - maximum exact KL:
0.4525210261
This is incremental V6-LoRA-versus-R2 KL, not total KL against the untouched official checkpoint. KL terms from different stages must not be added as though they were linear.
The KL result is disclosed as a divergence measurement; it is not presented as a standalone capability-retention score.
Checkpoint variants
- `main` branch: BF16 main/text checkpoint
- `f16` branch: F16 main/text checkpoint
- [GGUF repository](https://huggingface.co/0bserverx/RVN-Qwen3.8-Flash-Next-Abliterated-Uncensored-GGUF): 18 verified quant families, including the importance-matrix-built
IQ4_NL,IQ4_XS,IQ3_M,IQ3_S, andIQ3_XSfamilies - Coming separately: MLX variants and embedded-MTP builds
The HF checkpoints in this repository contain the main CausalLM only. The original vision tower and NextN/MTP speculative draft tensors are not included here. The linked GGUF builds are also text/main-only and exclude NextN/MTP. MLX and MTP artifacts will be published separately after their own load and generation verification.
These are very large sharded checkpoints. Check the repository inventory before downloading and plan CPU RAM, GPU memory, storage, and offload capacity accordingly.
Transformers usage
A recent Transformers runtime with Qwen4ExpForCausalLM support is required.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "0bserverx/RVN-Qwen3.8-Flash-Next-Abliterated-Uncensored"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
repo,
dtype=torch.bfloat16,
device_map="auto",
)For the F16 branch:
tokenizer = AutoTokenizer.from_pretrained(repo, revision="f16")
model = AutoModelForCausalLM.from_pretrained(
repo,
revision="f16",
dtype=torch.float16,
device_map="auto",
)Disable thinking for the evaluation-style answer surface:
rendered = tokenizer.apply_chat_template(
[{"role": "user", "content": "Your prompt"}],
tokenize=False,
add_generation_prompt=True,
enable_thinking=False,
)Research use, deployment responsibility, and disclaimer
This repository is intended for legitimate research, controlled evaluation, and lawful use, including interpretability research, alignment and refusal-behavior analysis, red-team testing, robustness work, and evaluation of model behavior under adversarial or sensitive prompts.
This is an abliterated, uncensored model—not a safety product, moderation service, compliance system, or ready-made production safety layer. It may generate inaccurate, offensive, explicit, unsafe, unlawful, or otherwise harmful material; follow dangerous instructions; reproduce biases; disclose information supplied in its context; or behave unpredictably outside the disclosed evaluations. The disclosed evaluations are narrow measurements, not a general safety certification or guarantee.
If you run, deploy, fine-tune, quantize, redistribute, or expose the model to other users, you are the operator and are responsible for the resulting system and its outputs. Before deployment, conduct your own risk assessment and implement controls appropriate to your use case and jurisdiction. Depending on context, those controls may include authentication, authorization, age or role restrictions, rate limits, sandboxing, data-loss prevention, content moderation, human review, logging, monitoring, abuse reporting, incident response, and emergency shutdown procedures.
Do not rely on model output as verified fact or as legal, medical, financial, security, or other professional advice. Do not provide personal, confidential, regulated, or security-sensitive data unless you have an appropriate lawful basis and adequate technical safeguards. You are responsible for evaluating output accuracy, legality, provenance, intellectual-property implications, and fitness for your intended purpose before acting on or distributing it.
Use of these files is subject to the Qwen Community License 1.0, the terms of any applicable third-party components or services, and all applicable laws and regulations. Public availability does not constitute permission or endorsement for any particular use. You are solely responsible for prompts submitted, outputs generated, downstream actions taken, and compliance obligations in your environment.
The model and accompanying materials are provided “AS IS” and “AS AVAILABLE,” without warranties of any kind, express or implied, including warranties of accuracy, reliability, safety, merchantability, fitness for a particular purpose, non-infringement, or uninterrupted operation. To the maximum extent permitted by applicable law, the repository maintainers and contributors disclaim liability for claims, damages, losses, or other consequences arising from access to, use of, inability to use, deployment of, or reliance on the model or its outputs.
Generated outputs are not statements, opinions, endorsements, or representations of the repository maintainers, contributors, Qwen, Alibaba, upstream authors, or their affiliated organizations. This repository does not speak for upstream authors or create additional obligations on their behalf; upstream rights and obligations remain governed by their own licenses and terms.
This section is a release notice, not legal advice. Obtain qualified legal and compliance advice for your deployment and jurisdiction.
Evaluation notes and limitations
- The strict refusal rate is
2/100; this is not presented as a zero-refusal model. - One protected response entered a repetition collapse after its refusal and continued until the 400-token cutoff.
98/100RR100 generations reached the fixed 400-token evaluation cap. Labels were assigned from the complete stored 400-token responses; truncation alone was not treated as refusal or failure.- The fixed evaluation gates do not establish comprehensive safety or universal capability retention.
- The release is text/main-only: no vision projector and no NextN/MTP draft head.
qwen4_expis a new architecture. Use a compatible recent Transformers, vLLM, or llama.cpp build.
Reproducibility and provenance
- V6 adapter SHA-256:
7525e257431415d95f5e8981ae7a01bc78de4850d7b3c7cd74c06b9316d22738 - V6 final receipt SHA-256:
62512bf4bce1c315fa69380d57616c07883ce71e284f95508741b902ff960a5e - Frozen RR100 prompt pack SHA-256:
fc66fa4ee363fdf26d3790bff2823f3784437ef524a8ba1be92e01712c5a33ce - Full RR100 response artifact SHA-256:
a18b392c649515b7db16da6d8bcafc2d1789eeb97cfa995004789dedaf531e5a - Full-response RR100 adjudication SHA-256:
6f83f6d7e347aad347c5ec5f159c6957d1ac81d8ace36e275df540da489b681a
License
The original Qwen Community License 1.0 applies. See LICENSE.
