CoolFace
Modelpublic

dhanesh-hf/Jarvis-Titan-V15-MoE-Merged

sourceHugging Faceotherupdated 11d agoView on Hugging Face
0likes196downloads
Model Card

J.A.R.V.I.S. Titan 14.8B MoE — Milestone M4 Standalone Merged Model

Official 100% standalone merged weights combining J.A.R.V.I.S. Titan 14.8B DeepSeekMoE with the Milestone M4 Calibrated Tri-Brid Adapter.

Architecture & Specifications

  • —Backbone: DeepSeekMoE 14.8B (1 Shared Expert + 8 Routed Experts, Top-2 active routing).
  • —Tri-Brid Strategic Bridge Layers: Layers [3, 7, 11, 15, 19, 23, 27] (7 memory bridge checkpoints).
  • —Tier 1 (Local SWA): Sliding Window Attention ($W = 2048$) with zero-copy GQA.
  • —Tier 2 (Salient Reservoir): Exact KV retrieval subspace ($D = 512$).
  • —Tier 3 (Titans Neural Memory): Bounded test-time learning recurrence ($\eta=10^{-3}$, $\rho=10^{-4}$, $\mu=0.95$, $||M||_F \le 50.0$).
  • —MAG-3 Adaptive Gating: Calibrated routing distribution targeting [$L$: ~55%, $R$: ~25%, $M$: ~20%].
  • —100% Zero-Loss Passthrough: Unbroken backbone residual stream for stable autoregressive generation.