CoolFace
Modelpublic

openeurollm/oellm-9b-256k-theta64m-prelude-anneal300b

sourceHugging Faceapache-2.0updated 6d agoView on Hugging Face
0likes206downloads
Model Card

OELLM 9B — 256K context — Prelude anneal300b

This is the OpenEuroLLM organization mirror of `birgermoell/oellm-9b-256k-theta64m-prelude-anneal300b`, pinned at source commit 4aad6daf4e9f472c491f2165bccc3db548cb3c31. Model weights, tokenizer, configuration, and evaluation records are unchanged.

This is a BF16 Hugging Face export of a 256K continued-pretraining checkpoint derived from `openeurollm/prelude`, revision f289699b246dba59907df27e743d4a433613175a (anneal300b_iter_0989075). It uses the Qwen3-compatible dense 9B architecture, a 262,144-token maximum position length, and RoPE theta 64,000,000.

This is a base completion model. It is not instruction-tuned or safety-aligned.

Relationship to the earlier release

The earlier `birgermoell/oellm-9b-256k-theta64m-prelude` uses the approximately 1T-token Prelude base. This repository instead starts from the approximately 300B-token annealed checkpoint. It is published separately as an experimental comparison lineage, not as an overwrite or drop-in quality replacement.

Architecture

PropertyValue
ArchitectureQwen3-compatible dense decoder-only transformer
Parametersapproximately 9B
Transformer layers36
Hidden / FFN size4,096 / 12,288
Attention heads / KV groups32 / 8 (GQA)
Head dimension128
Vocabulary262,144 tokens
NormalizationRMSNorm with Q/K layer normalization
MLPSwiGLU
Embeddingsuntied input/output embeddings
Maximum positions262,144
RoPE theta64,000,000
Weight dtypeBF16

The export preserves the Q/K normalization parameters, tokenizer, and positional configuration used in Megatron-LM.

Training lineage

StageSequence lengthRoPE thetaCPIterationsTokensLUMI job
16K16,384500,0001953999,292,92821683851
32K32,7681,000,0001476998,244,35221683852
64K65,5362,000,00027152,998,927,36021740962
128K131,07232,000,00082381,996,488,70421683854
256K262,14464,000,0001659989,855,74421684001

Total continued-pretraining volume: 7,982,809,088 tokens.

text
Prelude anneal300b
  -> 16K  (theta 500K)
  -> 32K  (theta 1M)
  -> 64K  (theta 2M)
  -> 128K (theta 32M)
  -> 256K (theta 64M)

The large final theta is deliberate: the staged ABF schedule keeps RoPE dimensions usable at retrieval distances far beyond the original short-context training range.

Optimization and systems configuration

  • —16 LUMI nodes / 128 AMD MI250X GCDs at every production stage.
  • —Tensor parallelism 8, pipeline parallelism 1, sequence parallelism enabled.
  • —Context parallelism increased to 16 at 256K.
  • —Micro-batch size 1 and global batch size 64.
  • —BF16, FlashAttention, distributed Adam (beta1 0.9, beta2 0.95), weight decay 0.1.
  • —Cosine learning-rate decay; 256K used 8e-6 to 8e-7 with approximately 5% warmup.
  • —Gradient clipping at 1.0 and selective activation recomputation.
  • —Megatron-Core torch_dist checkpoints; optimizer and RNG state are not part of this HF release.

The 256K production stage completed 59/59 iterations on launcher attempt 2, trained on 989,855,744 tokens, saved the final checkpoint successfully, and recorded zero skipped and zero NaN iterations. Attempt 1 ended during cold startup with a context-parallel NCCL timeout; the fresh second initialization then ran the entire stage stably. The last periodic training log at iteration 58 reported loss 1.427106. This is a run-health value, not a comparable benchmark score.

Continued-pretraining data

The 16K–128K curriculum used a frozen token-proportional multilingual blend with 152 prefixes. The 256K stage used a related 157-prefix blend targeted at very long sequences. Sources include FinePDFs, DCLM, HPLT3, multilingual synthetic text, Nemotron material, MegaMath, StarCoder, PeS2o, arXiv, and Wikipedia, grouped into length-aware tiers. The mix covers European languages plus code and scientific text, but is not uniform by language.

Conversion and release gates

The iteration-59 Megatron checkpoint is converted directly to a Qwen3-compatible Transformers layout in BF16. Publication requires the exact iteration marker, distributed metadata, non-empty weights and tokenizer, and a config with model_type=qwen3, max_position_embeddings=262144, and rope_theta=64000000. The Hub upload is verified for public visibility and required release files.

Quick multi-length retrieval evaluation

This is a deliberately small forced-choice needle-in-a-haystack evaluation for a base model. It tests English and Swedish at nominal 16K, 32K, 64K, 128K, and 256K lengths, with two trials at each of three depths. It verifies export and retrieval behavior but is not a stable benchmark estimate.

ContextEnglishSwedishCombined
16,384100.0% (6/6)100.0% (6/6)100.0% (12/12)
32,768100.0% (6/6)100.0% (6/6)100.0% (12/12)
65,536100.0% (6/6)100.0% (6/6)100.0% (12/12)
131,072100.0% (6/6)100.0% (6/6)100.0% (12/12)
262,144100.0% (6/6)100.0% (6/6)100.0% (12/12)

Depth 0.0 is the far beginning of the context (maximum query distance); depth 1.0 is nearest the query.

LanguageContextDepth 0.0Depth 0.5Depth 1.0
en16,384100.0% (2/2)100.0% (2/2)100.0% (2/2)
en32,768100.0% (2/2)100.0% (2/2)100.0% (2/2)
en65,536100.0% (2/2)100.0% (2/2)100.0% (2/2)
en131,072100.0% (2/2)100.0% (2/2)100.0% (2/2)
en262,144100.0% (2/2)100.0% (2/2)100.0% (2/2)
sv16,384100.0% (2/2)100.0% (2/2)100.0% (2/2)
sv32,768100.0% (2/2)100.0% (2/2)100.0% (2/2)
sv65,536100.0% (2/2)100.0% (2/2)100.0% (2/2)
sv131,072100.0% (2/2)100.0% (2/2)100.0% (2/2)
sv262,144100.0% (2/2)100.0% (2/2)100.0% (2/2)

Controls

LanguageControlAccuracy
enshort_ctx100.0% (2/2)
enshuffled100.0% (2/2)
enno_context50.0% (1/2)
svshort_ctx100.0% (2/2)
svshuffled100.0% (2/2)
svno_context0.0% (0/2)

Main-condition failures

  • —None in this small evaluation.

All individual records are available under evaluation/.

Evaluation method

The forced-choice base-LM harness inserts numeric key/value facts into a long context. The queried fact is placed at a selected depth, while the four candidate answers include adversarial distractor values that genuinely occur elsewhere in the same context. Candidates are ranked by answer-token log likelihood, avoiding any dependency on chat instruction following. Actual constructed contexts can be slightly shorter than the nominal boundary because complete fact lines are used and space is reserved for the query/candidate suffix.

RULER long-context evaluation

Results supplied by Jouni Luoma, using `NVIDIA/RULER`. Scores range from 0 to 100, with higher being better. Context-length columns are nominal lengths; — means that task was not run.

Single-needle retrieval

Task4K8K16K32K64K128K
niah_single_1100.00100.00100.00100.0099.8099.80
niah_single_2——————
niah_single_3——————

Multi-key, multi-value, and multi-query retrieval

Task4K8K16K32K64K128K
niah_multikey_1——————
niah_multikey_299.6096.8092.6088.8069.4041.40
niah_multikey_399.4088.4065.8051.4029.407.60
niah_multivalue99.0099.0598.3589.8083.0582.75
niah_multiquery——————

Variable tracking, common/frequent words, and question answering

Task4K8K16K32K64K128K
ruler_vt——————
ruler_cwe——————
ruler_fwe——————
ruler_qa_hotpot47.8047.6045.8044.4039.8038.60
ruler_qa_squad66.2252.6252.4550.5349.9044.95

No RULER results beyond 128K are included in this table.

Usage

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "openeurollm/oellm-9b-256k-theta64m-prelude-anneal300b"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

Keep rope_theta=64000000 and max_position_embeddings=262144 unchanged. Full-length inference requires substantial accelerator memory, memory-efficient attention, careful KV-cache sizing, and usually tensor parallelism.

Intended use

This checkpoint is intended for base-model and long-context research, continued pretraining, retrieval experiments, and as a starting point for later supervised or preference-based post-training. Use completion-style prompts; it is not a chat or instruction-following model.

Limitations

  • —The initial quick evaluation is intentionally small and covers only English and Swedish.
  • —Single-needle retrieval does not establish long-document reasoning, summarization, multi-needle retrieval, or robust generation throughout the entire 256K window.
  • —General and short-context regression suites remain necessary before broad quality claims.
  • —This 300B-base lineage should not be assumed to match the earlier 1T-base release.
  • —The model inherits limitations, biases, and uneven language representation from the Prelude base and continued-pretraining mixture.

Reproducibility

  • —Base revision: f289699b246dba59907df27e743d4a433613175a
  • —Megatron-LM commit: b359462c12858cedd2238a22eca0dca7aa6b8872
  • —Tokenizer SHA256: ac214401227404105817141c21c6693e3b17c878465a42765d2c345310e003ce
  • —16K–128K blend SHA256: c25d345c2c1969be4308fab2c5423584ef2fade72b1e711558fd459936b7ac6a
  • —256K blend SHA256: 18b8d441cbf0306a1a5935cec17f63a17d45ec9b71282675e594a3d1a1d51f3a
  • —256K checkpoint iteration: 59
  • —256K training job: 21684001

Related resources