linear-moe-hub/Gated-Deltanet-1.3B
8297
Model of the paper MoM: Linear Sequence Modeling with Mixture-of-Memories and Gated Delta Networks: Improving Mamba2 with Delta Rule.
The model was trained on a sample of SlimPajama with 100B tokens.
Model of the paper MoM: Linear Sequence Modeling with Mixture-of-Memories and Gated Delta Networks: Improving Mamba2 with Delta Rule.
The model was trained on a sample of SlimPajama with 100B tokens.