OsaurusAI/Ornith-1.5-9B-JANG_4D
<p align="center"><a href="https://osaurus.ai"><img src="./osaurus-x-banner.png" alt="Osaurus AI"></a></p>
OsaurusAI/Ornith-1.5-9B-JANG_4D
JANG_4D MLX bundle of ornith-ai/Ornith-1.5-9B — the recommended default — best size/quality balance.
Ornith 1.5 is an agentic coding / reasoning VLM built on a hybrid gated-delta linear attention + full attention backbone (3:1), with a 27-layer vision tower and native video support.
Bundle
How it was quantized
Three calibration methods, all driven by one capture pass — the per-input-channel second moment E[x_c^2] is simultaneously the Hessian diagonal, the imatrix weighting and the AWQ salient-channel statistic.
Tensors whose in_features is divisible by no MLX group size (the 27 vision linear_fc2 at 4304) stay fp16.
Modalities
Reasoning
Reasoning is ON by default — the no-kwarg generation prompt is byte-identical to enable_thinking=True and ends <|im_start|>assistant\n<think>\n.
It is toggleable, but note how: enable_thinking=False does not remove the think block, it prefills an empty closed one (<think>\n\n</think>\n\n). A parser testing merely for the presence of a <think> block will find one in both modes — test whether it has content.
There are no `reasoning_effort` tiers on this model family (unlike Qwen3.8). History <think> blocks are preserved unconditionally. Reasoning parser: qwen3; tool parser: qwen3_coder.
Sampling
Both presets from the vendor card are stamped into jang_config.json, and the coding preset is also written to generation_config.json so the two files agree.
Ornith 1.5 is an agentic coding model (SWE-bench Verified 79, Terminal-Bench 2.1 67.8), so this bundle defaults to the coding preset. Upstream's owngeneration_config.jsonships the general numbers (temp 1.0, presence 1.5) — usesampling_modes.generalif you want parity with the vLLM/Transformers defaults.
Stop tokens: [248046, 248044] (<|im_end|>, <|endoftext|>).
Speculative decoding (MTP)
The 9B checkpoint declares mtp_num_hidden_layers: 1 but ships no mtp.* weights, so there is no MTP head to preserve (mtp_mode: metadata_only_missing_weights). Native MTP exists on the 35B-A3B member of this family.
Credits
JANG quantization by Jinho Jang — <eric@osaurus.ai>
Base model: ornith-ai/Ornith-1.5-9B by Ornith AI.
