openeurollm/oellm-9b-256k-theta64m-prelude-anneal300b
OELLM 9B — 256K context — Prelude anneal300b
This is the OpenEuroLLM organization mirror of `birgermoell/oellm-9b-256k-theta64m-prelude-anneal300b`, pinned at source commit 4aad6daf4e9f472c491f2165bccc3db548cb3c31. Model weights, tokenizer, configuration, and evaluation records are unchanged.This is a BF16 Hugging Face export of a 256K continued-pretraining checkpoint derived from `openeurollm/prelude`, revision f289699b246dba59907df27e743d4a433613175a (anneal300b_iter_0989075). It uses the Qwen3-compatible dense 9B architecture, a 262,144-token maximum position length, and RoPE theta 64,000,000.
This is a base completion model. It is not instruction-tuned or safety-aligned.
Relationship to the earlier release
The earlier `birgermoell/oellm-9b-256k-theta64m-prelude` uses the approximately 1T-token Prelude base. This repository instead starts from the approximately 300B-token annealed checkpoint. It is published separately as an experimental comparison lineage, not as an overwrite or drop-in quality replacement.
Architecture
The export preserves the Q/K normalization parameters, tokenizer, and positional configuration used in Megatron-LM.
Training lineage
Total continued-pretraining volume: 7,982,809,088 tokens.
Prelude anneal300b
-> 16K (theta 500K)
-> 32K (theta 1M)
-> 64K (theta 2M)
-> 128K (theta 32M)
-> 256K (theta 64M)The large final theta is deliberate: the staged ABF schedule keeps RoPE dimensions usable at retrieval distances far beyond the original short-context training range.
Optimization and systems configuration
- 16 LUMI nodes / 128 AMD MI250X GCDs at every production stage.
- Tensor parallelism 8, pipeline parallelism 1, sequence parallelism enabled.
- Context parallelism increased to 16 at 256K.
- Micro-batch size 1 and global batch size 64.
- BF16, FlashAttention, distributed Adam (beta1 0.9, beta2 0.95), weight decay 0.1.
- Cosine learning-rate decay; 256K used
8e-6to8e-7with approximately 5% warmup. - Gradient clipping at 1.0 and selective activation recomputation.
- Megatron-Core
torch_distcheckpoints; optimizer and RNG state are not part of this HF release.
The 256K production stage completed 59/59 iterations on launcher attempt 2, trained on 989,855,744 tokens, saved the final checkpoint successfully, and recorded zero skipped and zero NaN iterations. Attempt 1 ended during cold startup with a context-parallel NCCL timeout; the fresh second initialization then ran the entire stage stably. The last periodic training log at iteration 58 reported loss 1.427106. This is a run-health value, not a comparable benchmark score.
Continued-pretraining data
The 16K–128K curriculum used a frozen token-proportional multilingual blend with 152 prefixes. The 256K stage used a related 157-prefix blend targeted at very long sequences. Sources include FinePDFs, DCLM, HPLT3, multilingual synthetic text, Nemotron material, MegaMath, StarCoder, PeS2o, arXiv, and Wikipedia, grouped into length-aware tiers. The mix covers European languages plus code and scientific text, but is not uniform by language.
Conversion and release gates
The iteration-59 Megatron checkpoint is converted directly to a Qwen3-compatible Transformers layout in BF16. Publication requires the exact iteration marker, distributed metadata, non-empty weights and tokenizer, and a config with model_type=qwen3, max_position_embeddings=262144, and rope_theta=64000000. The Hub upload is verified for public visibility and required release files.
Quick multi-length retrieval evaluation
This is a deliberately small forced-choice needle-in-a-haystack evaluation for a base model. It tests English and Swedish at nominal 16K, 32K, 64K, 128K, and 256K lengths, with two trials at each of three depths. It verifies export and retrieval behavior but is not a stable benchmark estimate.
Depth 0.0 is the far beginning of the context (maximum query distance); depth 1.0 is nearest the query.
Controls
Main-condition failures
- None in this small evaluation.
All individual records are available under evaluation/.
Evaluation method
The forced-choice base-LM harness inserts numeric key/value facts into a long context. The queried fact is placed at a selected depth, while the four candidate answers include adversarial distractor values that genuinely occur elsewhere in the same context. Candidates are ranked by answer-token log likelihood, avoiding any dependency on chat instruction following. Actual constructed contexts can be slightly shorter than the nominal boundary because complete fact lines are used and space is reserved for the query/candidate suffix.
RULER long-context evaluation
Results supplied by Jouni Luoma, using `NVIDIA/RULER`. Scores range from 0 to 100, with higher being better. Context-length columns are nominal lengths; — means that task was not run.
Single-needle retrieval
Multi-key, multi-value, and multi-query retrieval
Variable tracking, common/frequent words, and question answering
No RULER results beyond 128K are included in this table.
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "openeurollm/oellm-9b-256k-theta64m-prelude-anneal300b"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(
repo_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)Keep rope_theta=64000000 and max_position_embeddings=262144 unchanged. Full-length inference requires substantial accelerator memory, memory-efficient attention, careful KV-cache sizing, and usually tensor parallelism.
Intended use
This checkpoint is intended for base-model and long-context research, continued pretraining, retrieval experiments, and as a starting point for later supervised or preference-based post-training. Use completion-style prompts; it is not a chat or instruction-following model.
Limitations
- The initial quick evaluation is intentionally small and covers only English and Swedish.
- Single-needle retrieval does not establish long-document reasoning, summarization, multi-needle retrieval, or robust generation throughout the entire 256K window.
- General and short-context regression suites remain necessary before broad quality claims.
- This 300B-base lineage should not be assumed to match the earlier 1T-base release.
- The model inherits limitations, biases, and uneven language representation from the Prelude base and continued-pretraining mixture.
Reproducibility
- Base revision:
f289699b246dba59907df27e743d4a433613175a - Megatron-LM commit:
b359462c12858cedd2238a22eca0dca7aa6b8872 - Tokenizer SHA256:
ac214401227404105817141c21c6693e3b17c878465a42765d2c345310e003ce - 16K–128K blend SHA256:
c25d345c2c1969be4308fab2c5423584ef2fade72b1e711558fd459936b7ac6a - 256K blend SHA256:
18b8d441cbf0306a1a5935cec17f63a17d45ec9b71282675e594a3d1a1d51f3a - 256K checkpoint iteration: 59
- 256K training job: 21684001
Related resources
- Earlier 1T-base 256K checkpoint: `birgermoell/oellm-9b-256k-theta64m-prelude`
- 128K checkpoint from this anneal300b lineage: `birgermoell/oellm-9b-128k-theta32m-prelude-anneal300b`
- Implementation and evaluation code: `BirgerMoell/openeuro-longctx-datamix`
