CoolFace
Modelpublic

mdouglas/granite-3.1-3b-a800m-base-bnb-4bit

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes20downloads
Model Card
[!IMPORTANT] This repository is an experimental quantized version of the original model `ibm-granite/granite-3.1-3b-a800m-base`. It requires development versions of transformers and bitsandbytes.

Quantization

The MLP expert parameters have been quantized in the NF4 format along with all nn.Linear modules except lm_head and router modules, using an experimental bnb_4bit_target_parameters configuration option.

Granite-3.1-3B-A800M-Base

Model Summary Granite-3.1-3B-A800M-Base is a decoder-only language model to support a variety of text-to-text generation tasks. It extends the context length of Granite-3.0-3B-A800M-Base from 4K to 128K

Model Architecture: Granite-3.1-3B-A800M-Base is based on a decoder-only sparse Mixture of Experts (MoE) transformer architecture. Core components of this architecture are: Fine-grained Experts, Dropless Token Routing, and Load Balancing Loss.

Model2B Dense8B Dense1B MoE3B MoE
Embedding size2048409610241536
Number of layers40402432
Attention head size641286464
Number of attention heads32321624
Number of KV heads8888
MLP hidden size819212800512512
MLP activationSwiGLUSwiGLUSwiGLUSwiGLU
Number of Experts——3240
MoE TopK——88
Initialization std0.10.10.10.1
Sequence Length4096409640964096
Position EmbeddingRoPERoPERoPERoPE
# Parameters2.5B8.1B1.3B3.3B
# Active Parameters2.5B8.1B400M800M
# Training tokens12T12T10T10T