CoolFace
Modelpublic

oktayd/Qwen3.6-35B-v2-MoE-Ablit-Heretic-Uncensor-Hermes-MTP-Vision-Ollama

sourceHugging Faceapache-2.0updated 21d agoView on Hugging Face
1likes3.4kdownloads
Model Card

<p align="center"> <a href="https://nowpayments.io/donation?apikey=f1399916-9760-4d45-9129-b972124b8186" target="blank" rel="noreferrer noopener"> <img src="https://nowpayments.io/images/embeds/donation-button-white.svg" alt="Cryptocurrency & Bitcoin donation button by NOWPayments" /> </a> </p>

Qwen3.6-35B v2

[image]

๐Ÿ‘๏ธ Understand images. ๐Ÿ’ป Build things. ๐ŸŽญ Make it personal.

Ollama imports, including the vision projector.

BF16 / FT ยท GGUF / Llama ยท Ollama ยท v2 collection

Jump to: Model ยท Setup ยท Reliable runtime ยท Hardware ยท Additional notes ยท License ยท Collection

<a id="meet"></a>

โœจ Meet the model

<details> <summary>โœจ What you can do with it</summary>

Meet your local AI companion for ideas, code, images and conversation. Qwen3.6-35B v2 by oktayd brings these interests together in one downloadable model. Choose the edition that fits your setup, give it a task and shape its style with your own instructions.

Focus
๐Ÿง  Learn & exploreAsk about science, work through a problem or get an explanation in everyday language.
๐Ÿ’ป Build & fixTry a website idea, draft code or work through a bug together.
๐Ÿ› ๏ธ Connect your toolsUse it in a tool-enabled app for structured requests and workflows. Your app supplies and executes the tools.
๐Ÿ‘๏ธ Bring an imageAsk about a screenshot, document, diagram or scene.
๐ŸŽญ Set the personalityExplore stories, roleplay and different conversational styles through your instructions.
๐Ÿ“ฆ Run it your wayThree editions and seven GGUF sizes, from compact experiments to higher-precision weights.

These are things to try, not promises of perfect results. It can make mistakes or repeat itself; check important answers. Adult-oriented training is included (18+).

</details>

<details> <summary>๐Ÿง  A team of experts inside one model</summary>

MoE means Mixture of Experts. Think of a team of specialists: for each token, a router chooses which expert networks should contribute. This model has about 35 billion parameters in total, with roughly 3 billion active per token; its configuration selects 8 of 256 routed experts.

That saves computation compared with activating every expert at once. It does not turn a 35B download into a 3B-sized model: the expert weights still need disk space and accessible RAM/VRAM. Quantized editions make local use more practical. The experts are learned networks, not separate installed apps or named profession-specific agents.

</details>

<details> <summary>๐Ÿ•ถ๏ธ Street knowledge. Business mind. Your style.</summary>

The personality direction is direct, sharp-witted and business-minded: a street-smart conversation partner with humor, creative confidence and room for disagreement. The training mix includes internet culture, slang, practical business topics and personality-oriented conversations. Think less formal textbook, more an opinionated partner for brainstorming, writing and exploring alternatives.

You set the tone: professional, casual, blunt or playful. The aim is personality without blind agreement. This is a style and training focus, not a guarantee of factual expertise or flawless judgment.

</details>

<a id="training-data"></a> <details> <summary>๐Ÿ“š Training data & attribution โ€” 34,000 record uses across three runs</summary>

The three runs used 6,000 + 12,000 + 16,000 record uses. A record use is not necessarily a unique example; selected samples were used rather than entire upstream datasets. Training loss and record counts are not benchmark accuracy.

Training areaIncluded material
Knowledge, STEM & reasoningSelected instructional, science, mathematics and research-oriented examples
Coding, tools & agentsCode, debugging, software-engineering, tool-use and agent-workflow examples
Business & everyday workFinance, strategy, marketing, operations and practical question-answer examples
Vision & documentsSelected visual question-answering, chart, document, OCR, UI and 3D-related examples
Style & conversationWriting, roleplay, personality and multilingual caption examples

Private, manually prepared material accounts for 3,070 broad adult-learning uses and 133 synthetic 3D multi-view adult-learning uses; neither package is distributed. The public-source pool contains selected subsets from datasets in the categories above, including agents/search, coding, science, finance, security, visual QA and captions. Source URLs identify attribution only; their licenses are not replaced by this modelโ€™s license. Some retained local adapters have no recoverable upstream URL and are not represented as guessed sources.

</details>

<a id="architecture"></a> <details> <summary>โš™๏ธ Architecture & validation โ€” technical specifications, training scope and limits</summary>

<details> <summary>Technical specifications & validation</summary>

  • โ€”Qwen3_5MoeForConditionalGeneration; approximately 35B total parameters, 256 experts, 8 selected per token (the previous A3B label).
  • โ€”Native vision-language architecture; separate projector required for GGUF image input. Text and elementary red/blue-image smoke tests were recorded. Video, 3D consistency and full desktop/browser agents are not certified by those tests.
  • โ€”Training covered selected knowledge, instruction/agent, coding, personality and visual data: 34,000 record uses across 6,000 + 12,000 + 16,000. Record uses are not unique records. Training loss is not benchmark accuracy.
  • โ€”All 19 source MTP tensors were preserved. MTP/speculative acceleration is not enabled or validated by these benchmark results. Backend support must be tested separately.
  • โ€”Tool-call formatting and end-of-answer behavior are diagnostic targets, not guaranteed features. Known schema mistakes, factual errors and repetition remain.

</details>

</details>

<a id="quantization"></a> <details> <summary>๐Ÿ“ฆ Choose one quantization โ€” 7 GGUF downloads and hardware notes</summary>

VariantWeight fileSize (GiB)Status
IQ1_MQwen3.6-35B-v2-IQ1_M.gguf8.51experimental
IQ2_MQwen3.6-35B-v2-IQ2_M.gguf11.70experimental
IQ4_XSQwen3.6-35B-v2-IQ4_XS.gguf17.86released; broad quant-specific quality untested
Q3KMQwen3.6-35B-v2-Q3_K_M.gguf15.99released; broad quant-specific quality untested
Q4KMQwen3.6-35B-v2-Q4_K_M.gguf20.22released; broad quant-specific quality untested
Q5KMQwen3.6-35B-v2-Q5_K_M.gguf23.61released; broad quant-specific quality untested
Q8_0Qwen3.6-35B-v2-Q8_0.gguf35.21released; broad quant-specific quality untested

Add mmproj-Qwen3.6-35B-v2-F16.gguf (about 0.84 GiB) for image input. File size is not runtime RAM/VRAM: KV cache, activations, projector, runtime and OS require extra memory.

Q4KM/Q5KM/Q80 use standard llama.cpp quantization. IQ4XS/Q3KM/IQ1M/IQ2M use a small mixed-text importance-matrix calibration, excluding benchmark prompts. All variants were created directly from BF16, not by requantizing a low-bit file. These are not Unsloth Dynamic or Bartowski-branded exports.

IQ1_M/IQ2_M are experimental. Uncalibrated tensors, including MTP, stay at Q80; some MoE experts had incomplete calibration observations. The labels are not uniform bit widths for every weight. Substantial quality loss is possible. An IQ1M smoke answer 7 + 5 = 12 was mathematically right but violated number-only formatting; the original strict result remains available.

</details>

<a id="start-on-your-device"></a>

๐Ÿš€ Start on your device

Choose your hardware, then expand Ubuntu/Linux or Windows/PowerShell inside it. You only need one edition and one quantization.

<a id="reliable-runtime"></a> <details> <summary>โš™๏ธ Reliable runtime</summary>

For ordinary chat and agent work, use non-thinking mode, a finite output limit, and one model/request on a memory-constrained laptop. The included importer preserves the selected weight and projector hashes. ollama ps shows the active CPU/GPU allocation.

</details>

<a id="hardware-quickstarts"></a> <details> <summary>๐Ÿ–ฅ๏ธ H200 / large server โ€” full BF16 with Transformers</summary>

For a machine with enough GPU memory. Validated on Ubuntu / NVIDIA H200 with Transformers 5.16.1. BF16 weights alone are about 70 GB, plus runtime memory. A Windows client does not provide the remote server's VRAM; this package does not fit a single RTX 5090 or an 8-GB laptop GPU.

<details> <summary>๐Ÿง Ubuntu / Linux โ€” installation</summary>

Install Python 3.11, virtual-environment support and a CUDA-enabled PyTorch build from the official PyTorch selector that matches your NVIDIA driver. On a fresh Ubuntu machine:

bash
sudo apt-get update
sudo apt-get install -y python3-venv python3-pip
mkdir -p qwen-v2-server
cd qwen-v2-server
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
# Install the CUDA-enabled PyTorch command from the official selector here.
python -m pip install transformers==5.16.1 accelerate pillow
python -c "import torch; print(torch.cuda.is_available()); print(torch.cuda.get_device_name(0))"

If the GPU check fails, fix the driver/PyTorch installation before downloading the model. Save the shared Python example below as qwen_chat.py, then run python qwen_chat.py. The first run downloads the BF16 weights.

</details>

<details> <summary>๐ŸชŸ Windows / PowerShell โ€” native installation</summary>

Native Windows โ€” WSL2 is not required. Install Python 3.11 and a compatible NVIDIA driver, then use PowerShell 7. This native installation route is separate from the measured H200 validation, which was performed on Ubuntu:

powershell
New-Item -ItemType Directory -Path .\qwen-v2-server -Force | Out-Null
Set-Location .\qwen-v2-server
py -3.11 -m venv .venv
.\.venv\Scripts\python.exe -m pip install --upgrade pip
# Run the Windows CUDA PyTorch install command from the official selector,
# using .\.venv\Scripts\python.exe -m pip instead of pip.
.\.venv\Scripts\python.exe -m pip install transformers==5.16.1 accelerate pillow
.\.venv\Scripts\python.exe -c "import torch; print(torch.cuda.is_available()); print(torch.cuda.get_device_name(0))"

Save the shared example below as qwen_chat.py, then execute .\.venv\Scripts\python.exe .\qwen_chat.py. Use only Windows GPU/driver combinations supported by PyTorch and with sufficient memory.

If the H200 belongs to a remote Ubuntu server, Windows is just the client. Connect using Windows OpenSSH (replace the example and use your actual port/key), then follow the Ubuntu instructions in that remote shell:

powershell
ssh user@YOUR_SERVER

<details> <summary>๐Ÿง Optional: Ubuntu under WSL2</summary>

Optional alternative, not required for native Windows. Follow Microsoft's WSL installation guide. From an administrator PowerShell, only if you want this alternative:

powershell
wsl --install -d Ubuntu

Restart if requested and complete Ubuntu's first-run account setup. Then open Ubuntu:

powershell
wsl -d Ubuntu

Inside Ubuntu, follow this device's Ubuntu/Linux instructions and create a separate Linux Python environment. Keep the Windows NVIDIA driver current; use NVIDIA's CUDA-on-WSL guidance rather than installing a Linux display driver inside WSL. Verify GPU access before loading weights.

Use either the Windows or WSL model server during the test, not both on the same port. WSL does not add VRAM. Reuse model downloads where practical, but do not reuse Windows Python environments inside Linux; Ollama stores/imports may create another copy. Native Windows remains the default path above.

</details>

</details>

Shared Python example โ€” use after either installation:

python
import torch
from transformers import AutoProcessor, Qwen3_5MoeForConditionalGeneration

repo = "oktayd/Qwen3.6-35B-v2-MoE-Ablit-Heretic-Uncensor-Hermes-MTP-Vision-FT"
processor = AutoProcessor.from_pretrained(repo)
model = Qwen3_5MoeForConditionalGeneration.from_pretrained(
    repo, dtype=torch.bfloat16, device_map="auto"
)
messages = [{"role": "user", "content": [
    {"type": "text", "text": "Explain gravity briefly."}
]}]
inputs = processor.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True,
    return_dict=True, return_tensors="pt", enable_thinking=False
).to(model.device)
with torch.inference_mode():
    output = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(processor.batch_decode(
    output[:, inputs.input_ids.shape[1]:], skip_special_tokens=True
)[0])

For images, include an image content item supported by AutoProcessor. Large images and long contexts increase memory use.

</details>

<details> <summary>๐ŸŽฎ RTX 5090 ยท 32 GB โ€” Q4KM with llama.cpp</summary>

Use the GGUF / Llama edition. Q4KM is about 20.22 GiB, plus the 0.84-GiB projector and runtime memory. RTX 5090 / Q4KM was benchmarked with llama.cpp commit 427291b5b34cd914a31b3fd3b61a68f6184f4b9f on Ubuntu. The Windows steps below are installation guidance, not a separate Windows 5090 result.

<details> <summary>๐Ÿง Ubuntu / Linux โ€” installation & launch</summary>

Install Python, the HF CLI and a CUDA-enabled llama.cpp build. Put llama-server on PATH. Choose a driver/toolkit/build supporting the RTX 5090; a CPU-only binary is not sufficient for GPU offload.

bash
mkdir -p qwen-v2-q4
cd qwen-v2-q4
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade huggingface_hub
llama-server --list-devices
repo="oktayd/Qwen3.6-35B-v2-MoE-Ablit-Heretic-Uncensor-Hermes-MTP-Vision-Llama"
hf download "$repo" Qwen3.6-35B-v2-Q4_K_M.gguf mmproj-Qwen3.6-35B-v2-F16.gguf SHA256SUMS LICENSE --local-dir .
sha256sum --check --ignore-missing SHA256SUMS
llama-server -m Qwen3.6-35B-v2-Q4_K_M.gguf   --mmproj mmproj-Qwen3.6-35B-v2-F16.gguf   -ngl 99 -c 4096 -np 1 --host 127.0.0.1 --port 8080   --chat-template-kwargs '{"enable_thinking":false}'

Stop if checksums fail or the device listing does not show the intended NVIDIA GPU. In another terminal:

bash
curl http://127.0.0.1:8080/v1/chat/completions   -H "Content-Type: application/json"   -d '{"messages":[{"role":"user","content":"Explain gravity briefly."}],"max_tokens":256,"temperature":0,"stream":false}'

</details>

<details> <summary>๐ŸชŸ Windows / PowerShell โ€” native installation & launch</summary>

Install Python 3.11 and download the matching Windows CUDA package and any required CUDA runtime DLL package from the official llama.cpp releases. Extract the full package, preserving its DLLs; add that folder to PATH for this terminal. Do not substitute a CPU-only build. These commands assume llama-server.exe is on PATH. Use PowerShell 7 for native JSON argument handling.

powershell
New-Item -ItemType Directory -Path .\qwen-v2-q4 -Force | Out-Null
Set-Location .\qwen-v2-q4
py -3.11 -m venv .venv
.\.venv\Scripts\python.exe -m pip install --upgrade huggingface_hub
llama-server.exe --list-devices
$qwenRepo = "oktayd/Qwen3.6-35B-v2-MoE-Ablit-Heretic-Uncensor-Hermes-MTP-Vision-Llama"
.\.venv\Scripts\hf.exe download $qwenRepo Qwen3.6-35B-v2-Q4_K_M.gguf mmproj-Qwen3.6-35B-v2-F16.gguf SHA256SUMS LICENSE --local-dir .
if ($LASTEXITCODE -ne 0) { throw "Download failed" }
$qwenFiles = @("Qwen3.6-35B-v2-Q4_K_M.gguf", "mmproj-Qwen3.6-35B-v2-F16.gguf")
$qwenChecks = @{}
Get-Content .\SHA256SUMS | ForEach-Object {
    if ($_ -match '^([0-9a-fA-F]{64})  (.+)$') { $qwenChecks[$matches[2]] = $matches[1] }
}
foreach ($qwenFile in $qwenFiles) {
    if (-not $qwenChecks.ContainsKey($qwenFile)) { throw "Missing checksum: $qwenFile" }
    if ((Get-FileHash -LiteralPath $qwenFile -Algorithm SHA256).Hash -ne $qwenChecks[$qwenFile]) { throw "Checksum mismatch: $qwenFile" }
}
llama-server.exe -m Qwen3.6-35B-v2-Q4_K_M.gguf --mmproj mmproj-Qwen3.6-35B-v2-F16.gguf -ngl 99 -c 4096 -np 1 --host 127.0.0.1 --port 8080 --chat-template-kwargs '{"enable_thinking":false}'

Confirm that --list-devices lists the RTX 5090. In another PowerShell window:

powershell
$qwenBody = @{
    messages = @(@{ role = "user"; content = "Explain gravity briefly." })
    max_tokens = 256
    temperature = 0
    stream = $false
} | ConvertTo-Json -Depth 5
$qwenReply = Invoke-RestMethod -Uri "http://127.0.0.1:8080/v1/chat/completions" -Method Post -ContentType "application/json" -Body $qwenBody -TimeoutSec 300
$qwenReply.choices[0].message.content

<details> <summary>๐Ÿง Optional: Ubuntu under WSL2</summary>

Optional alternative, not required for native Windows. Follow Microsoft's WSL installation guide. From an administrator PowerShell, only if you want this alternative:

powershell
wsl --install -d Ubuntu

Restart if requested and complete Ubuntu's first-run account setup. Then open Ubuntu:

powershell
wsl -d Ubuntu

Inside Ubuntu, follow this device's Ubuntu/Linux instructions and create a separate Linux Python environment. Keep the Windows NVIDIA driver current; use NVIDIA's CUDA-on-WSL guidance rather than installing a Linux display driver inside WSL. Verify GPU access before loading weights.

Use either the Windows or WSL model server during the test, not both on the same port. WSL does not add VRAM. Reuse model downloads where practical, but do not reuse Windows Python environments inside Linux; Ollama stores/imports may create another copy. Native Windows remains the default path above.

</details>

</details>

Keep the service on localhost. For a remote server, use SSH port forwarding rather than opening an unauthenticated public endpoint. Start with a short context; other GPU workloads reduce available VRAM.

</details>

<details> <summary>๐Ÿ’ป Laptop ยท 32 GB RAM / 8 GB VRAM โ€” IQ4_XS with Ollama</summary>

IQ4_XS: about 17.86 GiB, plus a 0.84-GiB vision projector and runtime memory. It does not fit fully into 8 GB VRAM; GPU + CPU/RAM offload is required. Allow about 40 GB free disk for the download and a possible second imported copy. These are planning estimates, not measured peaks. Close memory-heavy apps and use one model at a time.

The complete IQ4_XS laptop comparison is still pending. The importer was previously tested with Ollama 0.33.3; installed versions and device-specific performance are recorded separately.

<details> <summary>๐Ÿง Ubuntu / Linux โ€” installation & launch</summary>

Install Ollama using its official Linux instructions, plus Python 3 and virtual-environment support. Start the Ollama service, or run ollama serve in a separate terminal if it is not already running. Do not start a second server on an occupied port.

bash
mkdir -p qwen-v2-iq4
cd qwen-v2-iq4
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade huggingface_hub requests
repo="oktayd/Qwen3.6-35B-v2-MoE-Ablit-Heretic-Uncensor-Hermes-MTP-Vision-Ollama"
hf download "$repo" Qwen3.6-35B-v2-IQ4_XS.gguf mmproj-Qwen3.6-35B-v2-F16.gguf import_ollama.py SHA256SUMS LICENSE --local-dir .
python import_ollama.py --quant IQ4_XS
ollama run qwen3.6-35b-v2:iq4_xs --think=false

</details>

<details> <summary>๐ŸชŸ Windows / PowerShell โ€” native installation & launch</summary>

Install Ollama for Windows and Python 3.11. Keep the Ollama app running. The commands use a dedicated Python environment and do not change PowerShell execution policy.

powershell
New-Item -ItemType Directory -Path .\qwen-v2-iq4 -Force | Out-Null
Set-Location .\qwen-v2-iq4
py -3.11 -m venv .venv
.\.venv\Scripts\python.exe -m pip install --upgrade huggingface_hub requests
$qwenRepo = "oktayd/Qwen3.6-35B-v2-MoE-Ablit-Heretic-Uncensor-Hermes-MTP-Vision-Ollama"
.\.venv\Scripts\hf.exe download $qwenRepo Qwen3.6-35B-v2-IQ4_XS.gguf mmproj-Qwen3.6-35B-v2-F16.gguf import_ollama.py SHA256SUMS LICENSE --local-dir .
if ($LASTEXITCODE -ne 0) { throw "Download failed" }
.\.venv\Scripts\python.exe .\import_ollama.py --quant IQ4_XS
if ($LASTEXITCODE -ne 0) { throw "Import failed" }
ollama run qwen3.6-35b-v2:iq4_xs --think=false

<details> <summary>๐Ÿง Optional: Ubuntu under WSL2</summary>

Optional alternative, not required for native Windows. Follow Microsoft's WSL installation guide. From an administrator PowerShell, only if you want this alternative:

powershell
wsl --install -d Ubuntu

Restart if requested and complete Ubuntu's first-run account setup. Then open Ubuntu:

powershell
wsl -d Ubuntu

Inside Ubuntu, follow this device's Ubuntu/Linux instructions and create a separate Linux Python environment. Keep the Windows NVIDIA driver current; use NVIDIA's CUDA-on-WSL guidance rather than installing a Linux display driver inside WSL. Verify GPU access before loading weights.

Use either the Windows or WSL model server during the test, not both on the same port. WSL does not add VRAM. Reuse model downloads where practical, but do not reuse Windows Python environments inside Linux; Ollama stores/imports may create another copy. Native Windows remains the default path above.

</details>

</details>

The importer verifies the selected weight/projector hashes and includes both. Check CPU/GPU allocation with ollama ps. No Transformers/BF16 download is needed. Ollama chat API.

</details>

<details> <summary>๐Ÿงช Smaller experimental option โ€” IQ2_M with Ollama</summary>

IQ2_M: about 11.70 GiB, plus a 0.84-GiB vision projector and runtime memory. It does not fit fully into 8 GB VRAM; GPU + CPU/RAM offload is required. Allow about 27 GB free disk for the download and a possible second imported copy. These are planning estimates, not measured peaks. Close memory-heavy apps and use one model at a time.

Experimental low-bit option. This is IQ2M, not Q2K. Reasoning, code and format accuracy can suffer. Initial small Windows smoke checks have started; a full laptop quality/performance comparison is still pending.

<details> <summary>๐Ÿง Ubuntu / Linux โ€” installation & launch</summary>

Install Ollama using its official Linux instructions, plus Python 3 and virtual-environment support. Start the Ollama service, or run ollama serve in a separate terminal if it is not already running. Do not start a second server on an occupied port.

bash
mkdir -p qwen-v2-iq2
cd qwen-v2-iq2
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade huggingface_hub requests
repo="oktayd/Qwen3.6-35B-v2-MoE-Ablit-Heretic-Uncensor-Hermes-MTP-Vision-Ollama"
hf download "$repo" Qwen3.6-35B-v2-IQ2_M.gguf mmproj-Qwen3.6-35B-v2-F16.gguf import_ollama.py SHA256SUMS LICENSE --local-dir .
python import_ollama.py --quant IQ2_M
ollama run qwen3.6-35b-v2:iq2_m --think=false

</details>

<details> <summary>๐ŸชŸ Windows / PowerShell โ€” native installation & launch</summary>

Install Ollama for Windows and Python 3.11. Keep the Ollama app running. The commands use a dedicated Python environment and do not change PowerShell execution policy.

powershell
New-Item -ItemType Directory -Path .\qwen-v2-iq2 -Force | Out-Null
Set-Location .\qwen-v2-iq2
py -3.11 -m venv .venv
.\.venv\Scripts\python.exe -m pip install --upgrade huggingface_hub requests
$qwenRepo = "oktayd/Qwen3.6-35B-v2-MoE-Ablit-Heretic-Uncensor-Hermes-MTP-Vision-Ollama"
.\.venv\Scripts\hf.exe download $qwenRepo Qwen3.6-35B-v2-IQ2_M.gguf mmproj-Qwen3.6-35B-v2-F16.gguf import_ollama.py SHA256SUMS LICENSE --local-dir .
if ($LASTEXITCODE -ne 0) { throw "Download failed" }
.\.venv\Scripts\python.exe .\import_ollama.py --quant IQ2_M
if ($LASTEXITCODE -ne 0) { throw "Import failed" }
ollama run qwen3.6-35b-v2:iq2_m --think=false

<details> <summary>๐Ÿง Optional: Ubuntu under WSL2</summary>

Optional alternative, not required for native Windows. Follow Microsoft's WSL installation guide. From an administrator PowerShell, only if you want this alternative:

powershell
wsl --install -d Ubuntu

Restart if requested and complete Ubuntu's first-run account setup. Then open Ubuntu:

powershell
wsl -d Ubuntu

Inside Ubuntu, follow this device's Ubuntu/Linux instructions and create a separate Linux Python environment. Keep the Windows NVIDIA driver current; use NVIDIA's CUDA-on-WSL guidance rather than installing a Linux display driver inside WSL. Verify GPU access before loading weights.

Use either the Windows or WSL model server during the test, not both on the same port. WSL does not add VRAM. Reuse model downloads where practical, but do not reuse Windows Python environments inside Linux; Ollama stores/imports may create another copy. Native Windows remains the default path above.

</details>

</details>

The importer verifies the selected weight/projector hashes and includes both. Its interactive defaults are 4,096 context tokens and 1,024 output tokens; the API examples use 2,048 / 256. Check CPU/GPU allocation with ollama ps. No Transformers/BF16 download is needed. Ollama chat API.

</details>

<a id="additional-notes"></a> <details> <summary>โž• Additional notes โ€” release integrity & benchmarks</summary>

<a id="release-integrity"></a> <details> <summary>๐Ÿงฐ Release integrity, verification & practical limits</summary>

SHA256SUMS verifies release files. RELEASE-NAMING.json maps old repository/file names to v2 and records unchanged weight hashes. Historical validation reports intentionally retain their original runtime names; the map resolves those names. RELEASE-SUMMARY.json and OLLAMA-VALIDATION.json record load, termination and elementary-answer checks separately.

Thinking mode can loop; start with thinking off and bounded output. Guard stops are incomplete answers, not successful corrections. The private training inputs, installer ZIP, raw benchmark prompts/answers and internal debug logs are not part of these end-user repositories. Keep use within applicable rights and deployment requirements.

</details>

<a id="benchmarks"></a> <details> <summary>๐Ÿ“Š Measured diagnostics โ€” scope matters</summary>

<details> <summary>Measured results, comparison settings & limitations</summary>

These are local Q4_K_M / llama.cpp diagnostics, not full official benchmark scores and not Ollama performance claims. Temperature 0, seed 42, thinking off, 16,384 context, 4,096 output cap, 45-second per-request budget. The RTX 2000 Ada used requested 24 GPU layers plus CPU offload; the actual offload-layer log was unavailable.

Model / deviceAttempted / 204CompletedOld strict pass / scoredEnd-to-end output tok/s (median, outputs โ‰ฅ64 tokens)
Q36 / H20020418661 / 108156.4
Q36 / RTX 509020419062 / 109174.6
Huihui / RTX 509020419687 / 114203.8
Q36 / RTX 2000 Ada (CPU+GPU)20419461 / 10918.2
Huihui / RTX 2000 Ada (CPU+GPU)18216976 / 9919.5

These strict counts omit pending official/manual evaluators, use changing denominators, and sometimes reject semantically correct formatting variants. Do not divide passes by all prompts or call these counts overall accuracy. A separate local regrade keeps content, protocol compliance and delivery apart. It does not certify unreviewed reasoning or execute generated code. The Huihui RTX 2000 Ada run left 22 tasks untested at the global deadline. Earlier BF16 diagnostics used thinking and are not directly comparable.

Huihui performed better on the shared automatically assessable RTX 5090 subset; no claim is made that this fine-tune universally surpasses its source or Qwen3.8. Rates mix generated lengths and are not pure hardware speedups. Native decode, prefill, client first-output and resource metrics are separated in BENCHMARK-DEVICE-SUMMARY.json. Raw prompts/answers and private training data are not published here.

Qwen3.8 comparison plan registers every benchmark family from the publisher card, including internal/unavailable tasks. Qwen3.8 has not been tested locally. Publisher scores use different harnesses, settings, annotations and trial counts and are shown only as references. No GPU job is launched by these support files.

</details>

<!-- q36-capability-expansion:start -->

๐Ÿ”Ž Broad benchmark โ€” tests and passes

Expand a device to inspect every test family. Passes / graded use the original strict evaluator, including format-sensitive checks. Ungraded answers are not failures or passes. Incomplete is a separate delivery flag and can overlap with ungraded. These are small local subsets, not official leaderboard scores.

<details> <summary>Q36 ยท H200 โ€” per-test breakdown</summary>

Test familyAttempted / plannedPasses / gradedUngradedIncompleteNot run
ARC-Challenge10 / 108 / 8220
BBH23 / 233 / 20330
ChartQA5 / 54 / 5000
GPQA-Diamond5 / 52 / 5000
GSM8K10 / 105 / 9110
HumanEval-Plus5 / 50 / 0500
IFEval10 / 100 / 01040
LiveCodeBench9 / 90 / 0900
MATH-Level-510 / 100 / 01020
MBPP-Plus5 / 50 / 0500
MMLU-Pro1 / 10 / 1000
MMMU6 / 64 / 6000
MathVista5 / 52 / 5000
Q36-Agent-Function-Calling5 / 50 / 5000
Q36-Answer-Termination-No-Looping10 / 107 / 10000
Q36-Benign-Compliance-No-Overrefusal20 / 200 / 02020
Q36-CAPTCHA-Detection-and-Handoff10 / 1010 / 10000
Q36-Contradiction-and-Anti-Sycophancy10 / 100 / 01020
Q36-Direct-Style-and-Personality10 / 100 / 01010
Q36-Hermes-Tool-Format5 / 55 / 5000
Q36-JSON-Schema5 / 50 / 5000
Q36-Legal-Alternatives-and-Boundaries10 / 100 / 01000
Q36-Output-Integrity5 / 55 / 5000
TruthfulQA10 / 106 / 9110

</details>

<details> <summary>Q36 ยท RTX 5090 โ€” per-test breakdown</summary>

Test familyAttempted / plannedPasses / gradedUngradedIncompleteNot run
ARC-Challenge10 / 108 / 8220
BBH23 / 233 / 20330
ChartQA5 / 54 / 5000
GPQA-Diamond5 / 51 / 5000
GSM8K10 / 107 / 10000
HumanEval-Plus5 / 50 / 0510
IFEval10 / 100 / 01020
LiveCodeBench9 / 90 / 0900
MATH-Level-510 / 100 / 01010
MBPP-Plus5 / 50 / 0500
MMLU-Pro1 / 10 / 1000
MMMU6 / 64 / 6000
MathVista5 / 52 / 5000
Q36-Agent-Function-Calling5 / 50 / 5000
Q36-Answer-Termination-No-Looping10 / 109 / 10000
Q36-Benign-Compliance-No-Overrefusal20 / 200 / 02020
Q36-CAPTCHA-Detection-and-Handoff10 / 1010 / 10000
Q36-Contradiction-and-Anti-Sycophancy10 / 100 / 01020
Q36-Direct-Style-and-Personality10 / 100 / 01000
Q36-Hermes-Tool-Format5 / 53 / 5000
Q36-JSON-Schema5 / 50 / 5000
Q36-Legal-Alternatives-and-Boundaries10 / 100 / 01000
Q36-Output-Integrity5 / 55 / 5000
TruthfulQA10 / 106 / 9110

</details>

<details> <summary>Huihui ยท RTX 5090 โ€” per-test breakdown</summary>

Test familyAttempted / plannedPasses / gradedUngradedIncompleteNot run
ARC-Challenge10 / 1010 / 10000
BBH23 / 2316 / 23000
ChartQA5 / 55 / 5000
GPQA-Diamond5 / 55 / 5000
GSM8K10 / 107 / 10000
HumanEval-Plus5 / 50 / 0500
IFEval10 / 100 / 01000
LiveCodeBench9 / 90 / 0950
MATH-Level-510 / 100 / 01020
MBPP-Plus5 / 50 / 0500
MMLU-Pro1 / 10 / 1000
MMMU6 / 64 / 6000
MathVista5 / 53 / 5000
Q36-Agent-Function-Calling5 / 50 / 5000
Q36-Answer-Termination-No-Looping10 / 1010 / 10000
Q36-Benign-Compliance-No-Overrefusal20 / 200 / 02000
Q36-CAPTCHA-Detection-and-Handoff10 / 1010 / 10000
Q36-Contradiction-and-Anti-Sycophancy10 / 100 / 01000
Q36-Direct-Style-and-Personality10 / 100 / 01000
Q36-Hermes-Tool-Format5 / 54 / 5000
Q36-JSON-Schema5 / 52 / 5000
Q36-Legal-Alternatives-and-Boundaries10 / 100 / 01000
Q36-Output-Integrity5 / 55 / 5000
TruthfulQA10 / 106 / 9110

</details>

<details> <summary>Q36 ยท RTX 2000 Ada (CPU+GPU) โ€” per-test breakdown</summary>

Test familyAttempted / plannedPasses / gradedUngradedIncompleteNot run
ARC-Challenge10 / 107 / 7330
BBH23 / 233 / 20330
ChartQA5 / 54 / 5000
GPQA-Diamond5 / 52 / 5000
GSM8K10 / 108 / 10000
HumanEval-Plus5 / 50 / 0500
IFEval10 / 100 / 01010
LiveCodeBench9 / 90 / 0900
MATH-Level-510 / 100 / 01010
MBPP-Plus5 / 50 / 0500
MMLU-Pro1 / 10 / 1000
MMMU6 / 61 / 6000
MathVista5 / 53 / 5000
Q36-Agent-Function-Calling5 / 50 / 5000
Q36-Answer-Termination-No-Looping10 / 108 / 10000
Q36-Benign-Compliance-No-Overrefusal20 / 200 / 02010
Q36-CAPTCHA-Detection-and-Handoff10 / 109 / 10000
Q36-Contradiction-and-Anti-Sycophancy10 / 100 / 01010
Q36-Direct-Style-and-Personality10 / 100 / 01000
Q36-Hermes-Tool-Format5 / 55 / 5000
Q36-JSON-Schema5 / 50 / 5000
Q36-Legal-Alternatives-and-Boundaries10 / 100 / 01000
Q36-Output-Integrity5 / 55 / 5000
TruthfulQA10 / 106 / 10000

</details>

<details> <summary>Huihui ยท RTX 2000 Ada (CPU+GPU) โ€” per-test breakdown</summary>

Test familyAttempted / plannedPasses / gradedUngradedIncompleteNot run
ARC-Challenge10 / 1010 / 10000
BBH11 / 236 / 101112
ChartQA5 / 55 / 5000
GPQA-Diamond5 / 53 / 3220
GSM8K10 / 107 / 10000
HumanEval-Plus5 / 50 / 0500
IFEval10 / 100 / 01000
LiveCodeBench9 / 90 / 0940
MATH-Level-510 / 100 / 01050
MBPP-Plus5 / 50 / 0500
MMLU-Pro1 / 10 / 0110
MMMU6 / 64 / 6000
MathVista5 / 53 / 5000
Q36-Agent-Function-Calling5 / 50 / 5000
Q36-Answer-Termination-No-Looping10 / 1010 / 10000
Q36-Benign-Compliance-No-Overrefusal10 / 200 / 010010
Q36-CAPTCHA-Detection-and-Handoff10 / 1010 / 10000
Q36-Contradiction-and-Anti-Sycophancy10 / 100 / 01000
Q36-Direct-Style-and-Personality10 / 100 / 01000
Q36-Hermes-Tool-Format5 / 53 / 5000
Q36-JSON-Schema5 / 53 / 5000
Q36-Legal-Alternatives-and-Boundaries10 / 100 / 01000
Q36-Output-Integrity5 / 55 / 5000
TruthfulQA10 / 107 / 10000

</details>

Machine-readable counts. Code generation is not a pass until the relevant execution tests have been graded.

๐Ÿ Hard benchmarks on the roadmap

All 25 families below are registered from the Qwen3.8-27B card. The matching official adapters and datasets are not yet fully prepared, and no local Qwen3.8 baseline has been run. Existing short similarly named diagnostics do not substitute for those runs.

<details> <summary>Coding, reasoning, agent and visual benchmarks โ€” complete planned list</summary>

BenchmarkPlanned local casesPreparation status
Terminal Bench 2.1 (Terminus)5Adapter/data preparation pending
SWE-bench Pro5Adapter/data preparation pending
NL2Repo-Bench5Adapter/data preparation pending
DeepSWE 1.15Adapter/data preparation pending
QwenSWEBenchTBDInternal release/access needed
CoWorkBenchTBDInternal release/access needed
JobBench5Adapter/data preparation pending
Agents' Last Exam5Adapter/data preparation pending
IFBench20Adapter/data preparation pending
GPQA Diamond20Adapter/data preparation pending
HLE10Adapter/data preparation pending
LiveCodeBench v610Adapter/data preparation pending
OSWorld-Verified5Adapter/data preparation pending
WebArena-Verified5Adapter/data preparation pending
AndroidWorld5Adapter/data preparation pending
RecreationBenchTBDInternal release/access needed
ClawEval-MM5Adapter/data preparation pending
SWE-MM5Adapter/data preparation pending
Vision2Web5Adapter/data preparation pending
MathVision10Adapter/data preparation pending
BabyVision10Adapter/data preparation pending
CharXiv (RQ)10Adapter/data preparation pending
OmniDocBench 1.510Adapter/data preparation pending
RealWorldQA20Adapter/data preparation pending
ERQA10Adapter/data preparation pending

</details>

Comparison graphics will follow measured results, with our local samples and publisher-reported scores kept clearly separate. Different harnesses, budgets and trial counts will not be presented as a head-to-head win. Full protocol and references.

</details>

</details>

๐Ÿ’ป Coding, agents and your laptop

The next local checks cover executable coding tests, bug fixes, tool calls, planning, recovery, memory and stopping at the right time. 64 case slots are specified, including 19 existing coding slots. They are compact skill diagnostics, not proof of AGI. Results and failures will both be reported; new capability claims require actual task-level evidence.

๐Ÿ› ๏ธ Laptop test plan, Hermes routing and Obsidian workflows. The runtime comparison and full memory integration are in preparation, not yet validated. The four device quickstarts above remain separate from these future end-to-end checks. <!-- q36-capability-expansion:end -->

<a id="license"></a>

๐Ÿ“œ License, lineage & credits

Apache-2.0 license. Dataset sources and their own license information are linked in DATASETS.md; model licensing does not relicense the source datasets. Thanks to the Qwen team, the inherited model and dataset authors, and the Soup, PEFT, Transformers, llama.cpp and Ollama projects.

Name guide: Qwen3.6-35B identifies the model family and approximate total parameter count; v2 is this project release; MoE means mixture of experts; Ablit, Heretic, Uncensor and Hermes describe inherited project lineage/training intent. MTP denotes preserved multi-token prediction weights, not a measured speedup; Vision denotes image-input support. FT, Llama and Ollama distinguish the three packages.

These names do not imply affiliation or a promise of unrestricted behavior. Full provenance and validation are retained, including the historical Opus4.7-labelled source. No universal superiority, guaranteed compliance or removal of memorization is claimed.

<details> <summary>๐Ÿงฌ What Ablit, Heretic, Uncensor, Hermes, MTP and Vision mean</summary>

This release builds on the previous project's model card, which records the following stages. These are inherited stages, not new operations performed while packaging this release.

Name / stageWhat it contributes
Qwen3.6 / MoEThe underlying language-and-vision architecture and mixture-of-experts backbone.
Reasoning-distilled lineage / Opus4.7The earlier card traces a lordx64 reasoning-distilled original followed by the huihui-ai derivative. The historical Opus4.7 label describes inherited reasoning-distillation provenance, not inclusion of proprietary Claude weights.
Ablit / abliteratedThe huihui-ai source underwent a weight-modification stage aimed at reducing refusal behavior. OBLITERATUS documents its MoE-capable nuclear method; this card does not reproduce the historical upstream run.
HereticA subsequent custom, fused-MoE-aware modification stage documented by the previous release. It is distinct from the earlier abliteration. OBLITERATUS documents the MoE-capable method; this card does not reproduce the historical upstream run.
OBLITERATUS NuclearA separately recorded inherited modification stage aimed at reducing refusal behavior. It is distinct from Ablit and Heretic. OBLITERATUS GitHub documents the project; the previous Q36 BF16 card records this lineage stage.
UncensorConcise v2 release-name label for the inherited refusal-reduction lineage; it does not guarantee unrestricted behavior.
HermesTool-oriented supervised training: function-call structure, coding, terminal/file/repository workflows and multi-tool coordination, using Hermes Function Calling and Hermes Agent reasoning traces. Your host application still supplies, authorizes and executes tools.
v2 fine-tuningThree additional Soup/PEFT training runs with 34,000 record uses spanning knowledge, instructions, coding, personality and selected vision data. See datasets and counts.
MTPPreserved multi-token-prediction tensors. Preservation is not evidence that speculative decoding is enabled or faster in your runtime.
VisionImage-input architecture; GGUF runtimes also need the matching projector. It does not itself supply browser control, memory or a 3D engine.

The old card reports 23,220 training and 1,179 validation examples for its own earlier SFT stage; those are separate from the current release's 34,000 record uses. It also records protection of 333 vision tensors and 19 MTP tensors in that earlier build. These historical checks are not new laptop benchmark results.

</details>

<a id="collection"></a>

๐Ÿ”— Related collection

The previous release remains separate: Qwen3.6 Opus4.7 Heretic Hermes Agent โ€” Editions.