abhaymin/Ornith-1.5-Cyclos-1M-Uncensored
Ornith 1.5 Cyclos 1M Uncensored — BF16
Native BF16 weights. A LoRA fine-tune of 0xKitkat/Ornith-1.5-35B-A3B-Uncensored, itself an uncensored derivative of [ornith-ai/Ornith-1.5-35B-A3B](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B). Merged into the base and extended to 1,048,576 context via YaRN.
Everything this model is good at comes from Ornith-1.5 — its coding and agentic post-training, its self-improvement RL, the native vision tower, the 262K context window and the multi-token-prediction head are all upstream work by the Ornith Team.
All credit for the underlying model belongs upstream — see Credits. This repository contributes a fine-tune, a context-window change, and nothing else.
For llama.cpp / Ollama / LM Studio, use the Q4_K_M GGUF build instead — it is far faster on consumer hardware. This BF16 repo exists for vLLM and SGLang, which cannot load that GGUF.
Changes from the base model
Stated per Apache-2.0 §4(b):
- LoRA fine-tune (SFT), merged into the base weights.
r=16,alpha=16,use_rslora=true(scalingalpha/sqrt(r)= 4.0). Targets attention and shared-expert projections plus, via PEFTtarget_parameters, the fused MoE expert tensorsmlp.experts.gate_up_projandmlp.experts.down_proj. Training focus: coding, market/trading analysis, content classification, and tool/function calling. - Context extended 262,144 -> 1,048,576 via YaRN in
config.json(rope_type: "yarn",factor: 4.0,original_max_position_embeddings: 262144).
The vision tower was not fine-tuned — every adapter tensor lives under language_model.*, and model.visual.* is bit-identical to the base after merging. Vision/video ability is inherited.
vLLM
Tested on 2 x RTX 5060 Ti (16 GB each, 31.7 GB total), CUDA 13.1, 59 GB RAM, vLLM 0.27.1. Loads and generates correctly.
vllm serve <this-repo> \
--tensor-parallel-size 2 \
--max-model-len 4096 \
--gpu-memory-utilization 0.92 \
--cpu-offload-gb 22 \
--enforce-eager--cpu-offload-gb is required on 32 GB of VRAM: the weights are ~67 GB, so roughly 44 GB spills to system RAM. Expect about 1.8 tok/s in that configuration — it is offload-bound, not compute-bound. Drop --cpu-offload-gb entirely once you have ~80 GB of VRAM (e.g. 2xA100 80GB, 4xL40S), and raise --max-model-len toward 1048576 as KV memory allows.
With enough VRAM, vLLM can also use the multi-token-prediction head for speculative decoding — it registers Qwen3_5MoeMTP, a head that llama.cpp discards:
vllm serve <this-repo> --tensor-parallel-size 4 --max-model-len 262144 \
--speculative-config '{"method":"mtp","num_speculative_tokens":1}'SGLang
The architecture is supported upstream (Qwen3_5MoeForConditionalGeneration); upstream Ornith recommends SGLang 0.5.9+. Not tested on this hardware — 32 GB of VRAM is below what a 67 GB BF16 model needs, and SGLang has no CPU-offload equivalent to vLLM's --cpu-offload-gb. Treat this as a starting point on a larger machine, not a verified command:
python -m sglang.launch_server \
--model-path <this-repo> \
--tp 2 \
--context-length 262144 \
--mem-fraction-static 0.90 \
--port 30000Transformers
from transformers import AutoModelForImageTextToText, AutoProcessor
import torch
model_id = "<this-repo>"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
model_id, dtype=torch.bfloat16, device_map="auto",
)Requires Transformers 5.8.1+ (Qwen3_5MoeForConditionalGeneration does not exist in 4.x).
The 1M context, in detail
config.json carries max_position_embeddings: 1048576 with:
"rope_parameters": { "rope_type": "yarn", "factor": 4.0,
"original_max_position_embeddings": 262144, ... }vLLM 0.27.1 reads the transformers-5 rope_parameters key (not the legacy rope_scaling), so the scaling applies automatically. --max-model-len still bounds what you actually allocate.
It is extrapolation, not training. The base is natively 262,144 with rope_type: default. YaRN interpolates RoPE frequencies at inference; nothing here was trained beyond 262,144, and no long-context evaluation was run. Treat 1M as an upper bound. For results you can trust, cap at the native length:
vllm serve <this-repo> --max-model-len 262144 ...Memory. Qwen3.5-MoE is hybrid — only 10 of 40 layers use full attention, the rest are linear/SSM with constant-size state — so KV is about 20 KiB/token at f16, roughly a quarter of a comparable dense model. At 1,048,576 that is still ~21 GiB of KV on top of ~67 GiB of BF16 weights, so full 1M in BF16 wants ~90 GiB of VRAM. On the 31.7 GB test rig this is not reachable; --max-model-len 4096 was used for verification.
If you are training on these weights, set the context back to native first. Training under YaRN while your sequences are a few thousand tokens applies the scaling to short positions, which is not how the base was trained:
"max_position_embeddings": 262144,
"rope_parameters": { "rope_type": "default", ... } // drop factor + original_max_position_embeddingsBuilt with soup-cli
Produced with [soup-cli](https://github.com/MakazhanAlpamys/Soup) — the LoRA merge, the YaRN context config, and the GGUF sibling build all ran through it. The rope block was generated by soup's own helper rather than hand-written:
from soup_cli.utils.long_context import get_rope_scaling_config
get_rope_scaling_config("yarn", 1_048_576, 262_144)
# {'type': 'yarn', 'factor': 4.0, 'original_max_position_embeddings': 262144}soup can serve and batch-run these weights directly, wrapping vLLM or SGLang:
soup serve -m <this-repo> --backend vllm --tp 2 --max-model-len 262144
soup serve -m <this-repo> --backend sglang --tp 2 --max-model-len 262144
soup infer -m <this-repo> -i prompts.jsonl -o out.jsonl
soup chat -m <this-repo>Note soup's backends are transformers, vllm, sglang and mii — there is no llama.cpp serving backend, so use llama-server directly for the GGUF sibling.
Honest limitations
- The 1M context is extrapolation, not training. The base is natively 262,144 with no RoPE scaling; YaRN extends it at inference. Quality degrades as you move beyond 262,144, and this has not been measured. Treat 1M as an upper bound, not a validated working length. To train on these weights, set
max_position_embeddingsback to 262144 andrope_typetodefaultfirst — training under YaRN on short sequences is not how the base was trained. - No audio. This architecture has no audio encoder;
<|audio_start|>-style tokens in the tokenizer are vestigial shared-vocabulary entries. Use an external ASR front-end. - No benchmarks are claimed. No evaluation suite was run. Upstream numbers do not transfer.
- "Uncensored" is inherited from the base, which reduced refusal behaviour. It does not mean every request is answered, and it does not shift responsibility for how the model is used.
Upstream recommended sampling: temperature=0.6, top_p=0.95, top_k=20. This is a reasoning model and emits <think>...</think> before its answer.
Credits and attribution
This is a derivative several steps down a chain. Credit belongs to, in order:
- [0xKitkat](https://huggingface.co/0xKitkat) — Ornith-1.5-35B-A3B-Uncensored, the direct base of this build (Apache-2.0).
- Ornith Team — ornith-ai/Ornith-1.5-35B-A3B, the coding/agentic post-training, vision tower, 262K context and MTP head (card declares MIT).
- Qwen Team — Qwen/Qwen3.6-35B-A3B (Apache-2.0).
- wangzhang / Abliterix — Qwen3.6-35B-A3B-abliterated, the abliterated donor and documented method.
- ggml-org — llama.cpp, the GGUF format, and the conversion/quantization tooling.
Licensed under Apache-2.0, inherited from the base model. Upstream components carry their own licenses as listed above.
