CoolFace
Modelpublic

zimplex/dllm-dreamreasoner-8b-finecode-approx-cf-v6-e0p1-r0p02-a0p1-b64-s24256-v1

sourceHugging Faceapache-2.0updated 2d agoView on Hugging Face
0likes736downloads
Model Card

DreamReasoner-8B FineCode approximate-CF-v6 CPT

This is the final Hugging Face-format checkpoint from continued pretraining of the released post-trained Dream-org/DreamReasoner-8B model on the sealed FineCode B64 dataset with the approximate-CF-v6 dependency graph.

Training configuration

  • —Run: dream-fc-cf-v6-e0p1-r0p02-a0p1-g64-m1-a1-16n-s24256-20260923-a1
  • —Base revision: ed62b1d2c82ccd234b05ed2463b4c0ee640f2068
  • —Source commit: 4f1ff9e7309fda97bd736b58fa0fc5b4151925fe
  • —Final optimizer step: 24,256
  • —Context length: 2,048
  • —Global batch: 64 sequences (microbatch 1, gradient accumulation 1)
  • —Hardware: 16 nodes x 4 GB300 GPUs
  • —Precision/distribution: BF16, FSDP full shard
  • —Learning rate: cosine decay from 5e-7 to 5e-8, warmup ratio 0.05
  • —Weight decay: 0.1
  • —Objective: projected_eq73_v1
  • —Graph: approximate CF v6, top-k 8
  • —Teacher parameters: epsilon=0.1, rollin_epsilon=0.02, action_epsilon=0.1
  • —Dependency release: zimplex/dllm-effect-parents-dreamreasoner-finecode-d1marginal-approx-w128-tau2-b64-v6
  • —Dependency manifest SHA-256: fd0222bf5a494349557507fdae66a457266b03cfd735b15c23bb2bd961906533
  • —Terminal-EOS supervision: enabled

The complete resolved training configuration is included as resolved_config.yaml.

Evaluation snapshot

Greedy B64/S64 confidence decoding at threshold 0.95 and seed 17:

BenchmarkGeneration capPass@1
HumanEval4,096147/164 (89.63%)
MBPP sanitized4,096210/257 (81.71%)
LiveCodeBench Pro4,096109/706 (15.44%)

Integrity

The original 14-file checkpoint payload (before this model card was added) has checkpoint fingerprint:

4f4c7eb135a70e8bbdd6466f8ebe1cc886c18bd45f65a49ca84f647336527e62

The source training-state manifest has SHA-256:

64869c62766db833611ef7a50ea7b3a198a784dccea9737e7e532eded36f5cf1

The weight file model.safetensors has SHA-256:

c6243cdffc4976de397aacbed660088278fcce624a592d36a92b00b303c02795

Loading

This repository contains custom DreamReasoner modeling code. Review it and use trust_remote_code=True when loading:

python
import torch
from transformers import AutoModel, AutoTokenizer

repo_id = "zimplex/dllm-dreamreasoner-8b-finecode-approx-cf-v6-e0p1-r0p02-a0p1-b64-s24256-v1"
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModel.from_pretrained(
    repo_id,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
)

This is a research checkpoint. It has not received a separate safety or deployment evaluation.