CoolFace
Modelpublic

CCSSNE/LuffyTheFox-Qwopus3.5-27B-v3-Uncensored-FernflowerAI-KL-ReLU-GGUF

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
0likes111downloads
Model Card

๐ŸŒŸ This is uncensored Qwopus3.5-27B-v3 made by Jackrong

๐ŸŒŸ Merged with Qwen3.5-27B-Uncensored-HauhauCS-Aggressive made by HauhauCS.

With repaired tensors, kullback leibler divergence ajdustment and ReLU finetunes.

Tensor repair by me. Method: Sig-ScaleSync-KL-ReLU (proprietary)

Quantization script available here with importance matrix support: https://pastebin.com/p6iN1f1Z

Repair Summary:

MetricValue
Total weight tensors546
Healthy538
C2-exempt (asymmetric, S<0.001)536
Repaired (C2)8
Skipped305
Time (pass1 / pass2)295.4s / 0.2s
Output size50.11 GB
RAM used0.65 GB

Repair Statistics Pass 1

Value
ฮฑ (min / mean / max)0.5202 / 0.5719 / 0.6609
D (min / mean / max)0.4141 / 0.5613 / 0.6536
S before โ†’ after0.0014 โ†’ 0.0005
Error reduction90.7%

Repaired Tensors Pass 1 (8)

TensorฮฑDS (before)S (after)
blk.60.ssm_conv1d.weight0.52020.6540.00170.0005
blk.52.ssm_conv1d.weight0.54590.6050.00150.0005
blk.56.ssm_conv1d.weight0.54970.5980.00150.0005
blk.53.ssm_conv1d.weight0.55200.5940.00150.0005
blk.57.ssm_conv1d.weight0.55640.5860.00150.0005
blk.62.ssm_conv1d.weight0.59130.5250.00130.0005
blk.58.ssm_conv1d.weight0.59840.5130.00130.0005
blk.61.ssm_conv1d.weight0.66090.4140.00110.0005

Repair Statistics Pass 2

Value
ฮฑ (min / mean / max)0.6172 / 1.2424 / 1.8899
D (min / mean / max)0.0050 / 0.3039 / 0.5591
S before โ†’ after0.0003 โ†’ 0.0004
KL before โ†’ after0.0322 โ†’ 0.0097
KL reduction69.9%
S error reduction74.4%

Repaired Tensors Pass 2 (16)

TensorฮฑDS (before)S (after)KL (before)KL (after)
blk.1.ssm_conv1d.weight1.43710.5590.00020.00030.08030.0180
blk.54.ssm_conv1d.weight0.61720.3940.00100.00040.05080.0127
blk.2.ssm_conv1d.weight1.41480.4890.00020.00030.04450.0064
blk.0.ssm_beta.weight1.88990.4350.00000.00010.04160.0012
blk.5.ssm_conv1d.weight1.41480.4440.00020.00040.03700.0085
blk.9.ssm_conv1d.weight1.33380.4020.00020.00040.03060.0043
blk.60.ssm_conv1d.weight1.02990.0050.00050.00050.02870.0241
blk.10.ssm_conv1d.weight1.36570.3470.00020.00040.02460.0034
blk.16.ssm_conv1d.weight1.31890.3950.00020.00040.02440.0047
blk.56.ssm_conv1d.weight0.92430.0050.00050.00040.02300.0176
blk.13.ssm_conv1d.weight1.28040.3860.00020.00030.02280.0062
blk.6.ssm_conv1d.weight1.36210.3280.00020.00040.02250.0029
blk.53.ssm_conv1d.weight1.02640.0050.00050.00050.02230.0174
blk.57.ssm_conv1d.weight0.89190.0050.00050.00040.02160.0182
blk.4.ssm_conv1d.weight1.26690.3470.00020.00040.02030.0074

๐ŸŒŸ Recommended Settings (LM Studio)

Chat template: pastebin.com/uk9ZkxCR (supports tool calling for Zed agent)

ParameterValue
Temperature0.7
Top K Sampling20
Presence Penalty1.5
Top P Sampling0.8
Min P Sampling0
Seed3407

System prompt: pastebin.com/pU25DVnB (solid) Or use this minimal string as the first line:

You are Qwen, created by Alibaba Cloud. You are a helpful assistant.

Then add anything you want after. Model may underperform without this first line.

๐Ÿ”ฅ Update (April 5): Iโ€™ve released the complete training notebook, codebase, and a comprehensive PDF guide to help beginners and enthusiasts understand and reproduce this model's fine-tuning process.

โค๏ธ Special thanks to the **Unsloth** open-source library and @KyleHessling1 for their support.

๐Ÿ“š Resources & Guides

๐Ÿ‘‰ [GitHub Repository: Jackrong-llm-finetuning-guide](https://github.com/R6410418/Jackrong-llm-finetuning-guide.git) Visit the repo to dive into the codebase and reproduce the results locally or on Colab.

๐Ÿ“ฅ Core Technical Document

๐Ÿ”— [Qwopus3.5-27b Complete Fine-Tuning Guide (PDF)](https://github.com/R6410418/Jackrong-llm-finetuning-guide/blob/main/guidePDF/Qwopus3-5-27b-Colab_complete_guide_to_llm_finetuning.pdf)

  • โ€”The Full Pipeline: A step-by-step walkthroughโ€”from downloading the base model and unifying heterogeneous data, to configuring trainer hyperparameters and publishing to Hugging Face.
  • โ€”Beginner Friendly: Includes an introductory guide to getting started with Google Colab and Unsloth.
A Note: My goal isn't just to detail a workflow, but to demystify LLM training. Beyond the social media hype, fine-tuning isn't an unattainable ritualโ€”often, all you need is a Google account, a standard laptop, and relentless curiosity. All training and testing for this project were self-funded. If you find this model or guide helpful, a Star โญ๏ธ on GitHub would be the greatest encouragement. Thank you! ๐Ÿ™
[!Note] The Claude series model optimizations are named under the Qwopus3.5 series, with the latest version being ๐ŸŒŸQwopus3.5-v3.

๐ŸŽฏ Motivation and Personal Opinion

HFXnLT4W4AA_2Ld

[!IMPORTANT] Qwopus 3.5-27B-v3 demonstrates exceptional token efficiency on LiveCodeBench, rapidly reaching high accuracy with minimal newly generated tokens. It also achieves the highest strict score of 95.73% (157/164) on the full HumanEval-164 benchmark under a conservative manual adjudication protocol, outperforming both the base Qwen3.5-27B and Claude-Distilled-v2. Thanks to Benjamin Marie (@bnjmn_marie) for sharing this chart and his insightful analysis!

Recent advances in language agents โ€” including systems such as OpenClaw โ€” have predominantly focused on improving reasoning accuracy through Chain-of-Thought (CoT) and self-reflection mechanisms, encouraging models to iteratively refine their reasoning before taking actions.

However, emerging evidence suggests that such "pre-action overthinking" is not always optimal for sequential decision-making. Instead, agent performance can be more effectively improved through a trial-and-error paradigm, where actions are executed early and refined based on environmental feedback.

๐Ÿ”ฌ Supporting Evidence

  • โ€”Reflexion[^1] demonstrates that agents can significantly improve decision-making by leveraging trial, error, and self-reflection โ€” shifting the role of reflection from pre-action deliberation to post-action correction, enabling agents to learn from concrete execution outcomes rather than speculative reasoning.
  • โ€”Post-failure reflection + retry[^2] substantially boosts performance:
  • โ€”๐Ÿ“ˆ +34.7% on mathematical reasoning tasks
  • โ€”๐Ÿ“ˆ +18.1% on function calling tasks

This provides strong empirical evidence that reflection is most effective when grounded in execution outcomes, rather than purely internal reasoning.

๐Ÿงญ My Approach

For multi-step and tool-augmented agent systems, performance should not be optimized solely through deeper pre-execution reasoning. A more effective strategy is an execution-driven optimization loop โ€” where agents perform lightweight initial reasoning, act in the environment, and iteratively refine their behavior based on feedback signals.

Paradigm Shift: from "reason-then-act" โ†’ "act-then-refine" The objective is not to achieve optimal reasoning in a single pass, but to enable robust task completion through iterative interaction and correction.

๐Ÿ’ก Model Introduction

Qwopus3.5-27B-v3 is a reasoning-enhanced model based on Qwen3.5-27B, designed to simultaneously improve reasoning stability and correctness while optimizing inference efficiency โ€” ultimately achieving stronger cross-task generalization capabilities, particularly in programming.

Key Highlights:

  • โ€”๐Ÿงฉ Structural Reasoning Optimization โ€” Refines the fundamental structure of the reasoning process through high-quality reasoning distillation and structural alignment, enabling higher accuracy rates via shorter, more stable reasoning paths.
  • โ€”๐Ÿ”ง Tool-Calling Reinforcement โ€” Incorporates specialized RL training for tool-calling, optimized for tool-augmented agent frameworks like OpenClaw, strengthening stability in continuous task execution and proficiency in tool invocation.
  • โ€”๐Ÿ” Act-Then-Refine Paradigm โ€” Designed for complex, multi-step agentic workflows, aligning with the core motivation of replacing pre-action deliberation with execution-driven refinement.

๐Ÿ”— Chain-of-Thought Optimization

๐Ÿšง The Problem with v2 Distillation

The v2 model was primarily trained through SFT on CoT data distilled from strong teacher models such as Claude. While this can transfer highโ€‘quality reasoning patterns, CoT traces from thirdโ€‘party datasets do not always faithfully reflect a modelโ€™s true internal reasoning process โ€” and after analysis, I found some portions may even be โ€œfabricatedโ€, meaning the traces were not actually generated by the claimed teacher model.[^3][^4]

Prior work further shows that CoT explanations can act as post-hoc rationalizations rather than genuine step-by-step reasoning[^3]. As a result, student models risk learning:

  • โ€”Surface-level pattern matching instead of underlying reasoning
  • โ€”Answer memorization rather than generalizable problem-solving
  • โ€”Reduced robustness on out-of-distribution tasks

โœ… What v3 Does Differently

v2 (Distillation)v3 (Structural Alignment)
CoT SourceThird-party distilled tracesCurated, verifiable reasoning chains
Learning TargetImitate teacher outputsLearn process-level reasoning
Reasoning StyleCompressed, potentially fabricatedExplicit, step-by-step, faithful
RobustnessLower on unseen tasksHigher generalization

v3 focuses on improving the faithfulness, completeness, and structural clarity of reasoning traces. Instead of imitating compressed teacher CoT, the model is trained to produce more explicit and verifiable intermediate steps โ€” enabling a transition from โ€œanswer imitationโ€ to process-level reasoning learning.

This improves both the interpretability and reliability of the reasoning process, providing a more stable foundation for downstream multi-step and agent-based tasks.

โš ๏ธ Side Effect: The generated CoT length in v3 will be slightly longer than v2, as a direct consequence of more explicit intermediate reasoning.

๐ŸŽ Qwopus3.5-27B-v3: Humaneval Benchmark Evaluation

๐Ÿ”ฌ Inference Setup: All models were evaluated under the Unsloth runtime using bfloat16 (BF16) precision โ€” optimally balanced for numerical range and memory efficiency at 27B scale. Answer verification, partial CoT adjudication, and statistical analysis were cross-validated by GPT-4.5-Pro (Thinking) and Claude Opus 4.6 (Thinking) to ensure reproducibility.

๐Ÿ“Š HumanEval โ€” 164-Task Full Benchmark

Three 27B-scale Qwen-family models were evaluated under a conservative manual adjudication protocol, addressing:

  • โ€”๐Ÿงน Code-extraction pollution
  • โ€”โœ‚๏ธ Answer / code separation issues
  • โ€”๐Ÿ—‚๏ธ Formatting noise in otherwise correct outputs
๐Ÿ† Result: Under this fair and strict evaluation setting, Qwopus3.5-27B-v3 achieves the best strict overall score of 95.73% (157/164) โ€” outperforming Qwen3.5-27B (94.51%, 155/164) and Claude-Distilled-v2 (92.68%, 152/164), while simultaneously reducing the number of manual rescues required.
ModelBase PassPlus Passvs. Qwen3.5-27B
๐Ÿฅ‡ Qwopus3.5-27B-v397.56% (160/164)95.73% (157/164)๐Ÿ“ˆ +1.22 pp
Qwen3.5-27B95.73% (157/164)94.51% (155/164)โ€” Baseline โ€”
Claude-Distilled-v295.12% (156/164)92.68% (152/164)๐Ÿ“‰ โˆ’1.83 pp

Screenshot 2026-04-01 at 11.25.34โ€ฏPM

Screenshot 2026-04-02 at 8.23.13โ€ฏAM


๐Ÿ—บ๏ธ Training Pipeline Overview

text
Base Model (Qwen3.5-27B)
 โ”‚
 โ–ผ
Qwen3.5-27B fine-tuned with Unsloth
 โ”‚
 โ–ผ
Supervised Fine-Tuning (SFT) + LoRA
(Response-Only Training masked on "<|im_start|>assistant\n<think>")
 โ”‚
 โ–ผ
Qwopus3.5-27B-v3

๐Ÿง  Example of Learned Reasoning Scaffold

The model includes targeted optimizations addressing Qwen3.5's tendency toward excessive or repetitive reasoning on simple queries. By distilling the structured reasoning habits of top-tier models like Claude Opus, Qwopus3.5-27B-v3 adopts a highly organized, step-by-step cognitive layout.

text
Example๏ผšThe user is asking about [Topic] and how it differs from [Topic B]. This is a [Task type] question. Let me break this down:

1. What is [Topic A]?
   - [Fact/Mechanism 1]
   - [Fact/Mechanism 2]
2. What is [Topic B]?
   - [Fact/Mechanism 1]
3. Key differences:
   - [Comparison Point 1]
   - [Comparison Point 2]

Let me make sure to be accurate: [...]
Actually, I should double-check: is [Fact] used before [Fact]? Yes, typically...
Let me provide a clear, well-structured answer:

๐Ÿ“š Training Data

The model was fine-tuned on a high-fidelity reasoning dataset, which was meticulously curated from a blend of premium open-source sources on Hugging Face. This dataset is the result of a rigorous mixing and cleaning process, specifically designed to filter out low-quality responses and ensure consistently strong logical performance across diverse analytical domains.

(Rest assured, the entire process is strictly by-the-book and 100% compliant with all terms and open-source licenses!)

โš ๏ธ Limitations & Intended Use

  • โ€”Hallucination Risk: While reasoning is strong, the model remains an autoregressive LLM; external facts provided during the thinking sequence may occasionally contain hallucinations if verifying real-world events.
  • โ€”Intended Scenario: Best suited for offline analytical tasks, coding, math, and heavy logic-dependent prompting where the user needs to transparently follow the AI's internal logic.
  • โ€”This model is a test version intended solely for learning and demonstration purposes, and is for academic research and technical exploration use only.
  • โ€”Developer Disclaimer: This is an independent, personal project. Since the developer lacks the specialized technical resources and infrastructure of a large-scale industrial lab, the model's reasoning chain (CoT) may occasionally exhibit instability, logic loops, or reasoning drift. Users are advised to use this model with these experimental limitations in mind.
Note: The test results presented here differ from the scores on the 27B-v2 model card because the context length was increased for this evaluation. Consequently, the number of tasks affected by context window truncation has changed for each model, leading to different final scores. Please ensure comparisons are made under the same variable settings.

All post-evaluation standard result files will be uploaded to this repository for transparency and reproducibility. These include:

  • โ€”Jackrong_Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2_humaneval_all_evalonly_eval_results
  • โ€”Jackrong_Qwopus3.5-27B-v3-test1_humaneval_all_evalonly_eval_results
  • โ€”qwen_Qwen3.5-27B_humaneval_all_evalonly_eval_results

โš ๏ธ Note on evaluation artifacts. The released result files are based on raw model generations, which may contain formatting issues (e.g., Markdown wrappers, answer/code mixing), truncation, or minor token-level corruption. As an independent project operating under limited resources, the evaluation scope here is intentionally focused rather than exhaustive โ€” a comprehensive, multi-domain assessment comparable to large institutional releases was not feasible. Capabilities beyond those benchmarked remain unverified, and users are encouraged to evaluate suitability against their own task requirements before adoption.

๐Ÿ™ Acknowledgements

Significant thanks to the Unsloth AI team for making rapid fine-tuning of large LLM models accessible. Additionally, we acknowledge Qwen internally, and the open-source community developers producing exceptional distilled datasets.

This qwen3_5 model was trained 2x faster with Unsloth and Huggingface's TRL library.

<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/>

References

[^1]: Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv:2303.11366.

[^2]: Bensal, S., Jamil, U., Bryant, C., Russak, M., Kamble, K., Mozolevskyi, D., Ali, M., & AlShikh, W. (2025). Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning. arXiv:2505.24726. https://arxiv.org/abs/2505.24726

[^3]: Anthropic (2025). Reasoning Models Don't Always Say What They Think. https://www.anthropic.com/research/reasoning-models-dont-say-think

[^4]: Lyu et al. (2023). Faithful Chain-of-Thought Reasoning. ACL. https://aclanthology.org/2023.ijcnlp-main.20/

๐Ÿ“– Citation

If you use this model in your research or projects, please cite:

bibtex
@misc{jackrong_qwen35_27b_v3
  title        = {Jackrong/Qwopus3.5-27B-v3},
  author       = {Jackrong},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/Jackrong/Qwopus3.5-27B-v3}}
}