JetBrains/Mellum2-12B-A2.5B-Base-Pretrain
<img alt="Mellum" src="mellum-logo-dark.svg" width="320">
Mellum2 Base Pretrain
[!Note] Use this checkpoint as a starting point for research on long-context extension or for 8K-context continued pretraining and fine-tuning. For downstream applications use Base, Instruct, or Thinking instead.
Mellum2 Base Highlights
Mellum2 Base is a pretrained causal language model trained by JetBrains.
The model uses a Mixture-of-Experts architecture with 64 experts and activates 8 experts per token. It uses a combination of sliding-window and full attention layers, with a context length of 8,192 tokens.
This is a checkpoint before long-context extension.
Mellum2 Model Family
This repository contains one checkpoint from the Mellum2 family.
Model Overview
Mellum2 Base has the following features:
- Number of Layers: 28
- Hidden Size: 2304
- Intermediate Size: 7168
- MoE Intermediate Size: 896
- Number of Experts: 64
- Number of Activated Experts: 8
- Number of Attention Heads (GQA): 32 for Q and 4 for KV
- Context Length: 8,192
- Sliding Window: 1,024
- Vocabulary Size: 98,304
- Precision: bfloat16
- License: Apache 2.0
Serving with vLLM
This checkpoint has an 8K context length (long-context extension is applied in Base).
vllm serve JetBrains/Mellum2-12B-A2.5B-Base-Pretrain --max-model-len 8192Quickstart
Text-Only Input (base model — use the completions endpoint, not chat)
from openai import OpenAI
# Configured by environment variables
client = OpenAI()
completion = client.completions.create(
model="JetBrains/Mellum2-12B-A2.5B-Base-Pretrain",
prompt="def fibonacci(n):\n ",
max_tokens=4096,
temperature=0.6,
top_p=0.95,
extra_body={
"top_k": 20,
},
)
print("Completion:", completion)Evaluation
Evaluation results are available in the model card. All values are self-reported by JetBrains.
For more details, see the Mellum2 Technical Report.
License
Released under the Apache 2.0 license.
