CoolFace
Modelpublic

leonsarmiento/Ornith-Agents-A1-3.6-35B-A3B-dare_ties-6bit-XL-mlx

sourceHugging Faceupdated 2mo agoView on Hugging Face
1likes135downloads
Model Card

leonsarmiento/Ornith-Agents-A1-3.6-35B-A3B-dare_ties-6bit-XL-mlx

This model was converted to MLX format from `tepirale/Ornith-Agents-A1-3.6-35B-A3B-dare_ties` using BaseQuant_XL 6/8-bit mixed quantization optimized for Apple Silicon. The vision encoder is preserved and quantized at 6-bit, making this a full multimodal model.

BaseQuant_XL keeps the most routing-critical layers in full bf16 precision — the MoE router gate, shared expert gate, shared expert, and lm_head — while applying aggressive quantization to the bulk parameters. This preserves routing accuracy and output quality where it matters most.

About XL Quantization

BaseQuant_XL is a fully data-agnostic, static quantization. No calibration dataset, no sensitivity analysis, no importance matrix. Precision is allocated purely by architectural role — routing-critical layers get higher precision, bulk expert parameters get lower precision. The result is a transparent, faithful capture of the source model.

Data-dependent calibration quantizations (iMatrix, AWQ, GPTQ, oQ, oQ4e, etc.) use a calibration set to guide bit allocation. This can produce a skewed representation of the model: domains well-represented in the calibration data (English, popular topics, public or leaked benchmarks) are preserved better, while underrepresented domains (non-English languages, niche use cases, your own data) are preserved worse. XL avoids this trade-off entirely — it generalizes honestly because it is never fit to any particular data distribution.

Model Description

This is a 50/50 DARE-TIES merge of two complementary Qwen3.5-35B-A3B agentic models:

Source ModelWeightDensityFocus
InternScience/Agents-A10.50.6General agentic abilities: long-horizon search, engineering, scientific research, instruction following, tool-calling
deepreinforce-ai/Ornith-1.0-35B0.50.6RL-tuned agentic coding (Terminal-Bench 64.2, SWE-bench Verified 75.6)
Qwen/Qwen3.5-35B-A3Bbase—Base model

The merge combines Ornith's coding and terminal-task strength with Agents-A1's broader tool-use and research capabilities. The architecture is a 35B Mixture-of-Experts model with only ~3B active parameters per token, featuring 256 experts (8 active + 1 shared), hybrid full + linear (Gated DeltaNet) attention, a vision encoder, and an extended 262K context window.

Use with mlx

bash
pip install -U mlx-vlm
bash
python -m mlx_vlm.generate --model leonsarmiento/Ornith-Agents-A1-3.6-35B-A3B-dare_ties-6bit-XL-mlx --max-tokens 256 --temperature 0.85 --top-p 0.95 --top-k 20 --min-p 0.01 --repeat-penalty 1.05 --prompt "Hello"

BaseQuant_XL Quantization Strategy

Bit DepthLayersRationale
bf16 (unquantized)mlp.gate (router), shared_expert_gate, lm_head, shared_expertRouting decisions and shared computation path — errors here are qualitatively different from precision loss
8-bitembed_tokens, self_attn (full attention), linear_attn (DeltaNet)Every-token layers with moderate sensitivity — 8-bit is near-lossless
6-bitvision_tower, switch_mlp (routed experts)Bulk of parameters, only 8 of 256 experts active per token — natural redundancy tolerates lower precision

Quantization Details

LayerBitsGroup Size
mlp.gate (router)bf16—
shared_expert_gatebf16—
lm_headbf16—
shared_expertbf16—
embed_tokens864
self_attn (full attention)864
linear_attn (DeltaNet)864
vision_tower664
switch_mlp (routed experts)664
Default fallback864
  • —Quantization type: BaseQuant_XL mixed (multimodal, vision preserved)
  • —Group size: 64
  • —Method: Custom quant_predicate via mlx_vlm

Recommended Inference Parameters

Inherited from Agents-A1 (more conservative, broader tool-use optimization):

ParameterValue
temperature0.85
top_p0.95
top_k20
min_p0.01
repeat_penalty1.0
presence_penalty1.1
Note: These use Agents-A1's recommended settings rather than averaging across both parents. Ornith's original values (temp 1.0, topp 1.0, topk 40) were optimized for Terminal-Bench specifically — Agents-A1's broader agentic profile is a safer default for general use. Adjust as needed.

Reasoning and Tool-Call Parsing

ParserValue
reasoning_parserqwen3
tool_call_parserqwen3_coder
This is a Qwen3.5-based model — preserve_thinking is not applicable.

Benchmarks (n=30, 5-bit XL MLX)

Benchmark comparison of both merges against parent models and the Qwen3.6-35B-A3B base. Higher bit depth (6-bit) is expected to match or slightly exceed these results:

[image]

BenchmarkAgents-A1 (parent)Ornith-1.0 (parent)**DARE-TIES merge**Task Arithmetic mergeQwen3.5-35B-A3B (base)
MMLU70.063.370.073.356.7
MMLU_PRO53.360.056.750.056.7
HELLASWAG86.783.386.783.383.3
TRUTHFULQA100.0100.096.793.3100.0
ARC_CHALLENGE90.083.390.090.086.7
WINOGRANDE76.776.773.376.780.0
HUMANEVAL86.766.786.786.760.0
MBPP80.076.780.076.786.7
MATHQA (thinking)96.796.793.396.793.3
LIVECODEBENCH40.043.343.336.740.0

The DARE-TIES merge inherits Agents-A1's strength on HUMANEVAL, MMLU, and ARCCHALLENGE while partially recovering Ornith's MMLUPRO edge. Minor regressions on TRUTHFULQA, WINOGRANDE, and MATHQA compared to parents — a typical DARE-TIES trade-off. Both merges clearly dominate the Qwen3.6 base on most knowledge and coding benchmarks.