zimplex/dllm-dreamreasoner-8b-finecode-approx-cf-v6-e0p1-r0p02-a0p1-b64-s24256-v1
DreamReasoner-8B FineCode approximate-CF-v6 CPT
This is the final Hugging Face-format checkpoint from continued pretraining of the released post-trained Dream-org/DreamReasoner-8B model on the sealed FineCode B64 dataset with the approximate-CF-v6 dependency graph.
Training configuration
- Run:
dream-fc-cf-v6-e0p1-r0p02-a0p1-g64-m1-a1-16n-s24256-20260923-a1 - Base revision:
ed62b1d2c82ccd234b05ed2463b4c0ee640f2068 - Source commit:
4f1ff9e7309fda97bd736b58fa0fc5b4151925fe - Final optimizer step: 24,256
- Context length: 2,048
- Global batch: 64 sequences (microbatch 1, gradient accumulation 1)
- Hardware: 16 nodes x 4 GB300 GPUs
- Precision/distribution: BF16, FSDP full shard
- Learning rate: cosine decay from
5e-7to5e-8, warmup ratio 0.05 - Weight decay: 0.1
- Objective:
projected_eq73_v1 - Graph: approximate CF v6, top-k 8
- Teacher parameters:
epsilon=0.1,rollin_epsilon=0.02,action_epsilon=0.1 - Dependency release:
zimplex/dllm-effect-parents-dreamreasoner-finecode-d1marginal-approx-w128-tau2-b64-v6 - Dependency manifest SHA-256:
fd0222bf5a494349557507fdae66a457266b03cfd735b15c23bb2bd961906533 - Terminal-EOS supervision: enabled
The complete resolved training configuration is included as resolved_config.yaml.
Evaluation snapshot
Greedy B64/S64 confidence decoding at threshold 0.95 and seed 17:
Integrity
The original 14-file checkpoint payload (before this model card was added) has checkpoint fingerprint:
4f4c7eb135a70e8bbdd6466f8ebe1cc886c18bd45f65a49ca84f647336527e62
The source training-state manifest has SHA-256:
64869c62766db833611ef7a50ea7b3a198a784dccea9737e7e532eded36f5cf1
The weight file model.safetensors has SHA-256:
c6243cdffc4976de397aacbed660088278fcce624a592d36a92b00b303c02795
Loading
This repository contains custom DreamReasoner modeling code. Review it and use trust_remote_code=True when loading:
import torch
from transformers import AutoModel, AutoTokenizer
repo_id = "zimplex/dllm-dreamreasoner-8b-finecode-approx-cf-v6-e0p1-r0p02-a0p1-b64-s24256-v1"
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModel.from_pretrained(
repo_id,
trust_remote_code=True,
torch_dtype=torch.bfloat16,
)This is a research checkpoint. It has not received a separate safety or deployment evaluation.
