CoolFace
Modelpublic

violetxi/qwen35-9b-harvey-v4-notes-conditioned-100m-kl-0p05

sourceHugging Faceapache-2.0updated 4d agoView on Hugging Face
0likes30downloads
Model Card

qwen35-9b-harvey-v4-notes-conditioned-100m-kl-0p05

Full Qwen3.5-9B model trained on 100M notes + note-conditioned trajectory tokens, with KL coefficient 0.05. This revision is checkpoint-3068, after 2 epoch(s) and 3,068 updates.

RevisionCheckpointEpoch
main, final, checkpoint-3068checkpoint-30682
epoch1, checkpoint-1534checkpoint-15341

The dataset contains 99,998,917 supervised tokens per epoch before causal shifting: 69,999,985 note tokens and 29,998,932 assistant trajectory tokens. After shifting, each epoch has 99,993,091 supervised tokens. The full run trains for two epochs; KL replay is additional.

Objective: notes next-token prediction + assistant-masked note-conditioned trajectory next-token prediction + 0.05 × KL(base || student). The frozen reference is Qwen/Qwen3.5-9B at c202236235762e1c871ad0ccb60c8ee5ba337b9a. Replay comes from harvey-kl-ground-sessions: 2,048 sessions, context limit 8,192, up to 128 assistant prediction positions per draw. KL uses the full vocabulary and the mean of session means.

Training uses context rows of 16,384 tokens, global batch 8, gradient accumulation 1, learning rate 5e-6 with cosine schedule, warmup ratio 0.03, seed 0, and 8 A100. Slurm training job: 192149. Online W&B run.

Held-out KL at this checkpoint: 0.010288765895, over 256 sessions / 32,768 prediction positions. This measures drift from the base. Closed-book recall evaluation is complete; the native open-book attempt is incomplete. See the evaluation artifacts below.

Four composite safetensors shards contain the trained text weights converted to the pinned base serving dtype, with the base auxiliary components retained. Config, tokenizer, chat template, processors, and generation settings are included. No adapter or local reconstruction is required.

python
from transformers import AutoModelForImageTextToText, AutoTokenizer
model_id = "violetxi/qwen35-9b-harvey-v4-notes-conditioned-100m-kl-0p05"
revision = "main"  # main: final; epoch1: first epoch
tokenizer = AutoTokenizer.from_pretrained(model_id, revision=revision)
model = AutoModelForImageTextToText.from_pretrained(
    model_id, revision=revision, dtype="auto", device_map="auto"
)

Evaluation artifacts — 2026-09-23

Closed-book recall probes: complete, 7,933 probes. Evaluation dataset: metrics and per-probe outcomes. Scoring is deterministic and uses no GPT judge.

Native open-book: incomplete; no full benchmark score. Requested 250 tasks × four samples, 20 turns, GPT-5.6 Sol grading. The job failed on a model-generation API timeout. Incomplete-run status. Diagnostic preflight grades are excluded from benchmark scores.

Native grading is preserved. A known-correct diagnostic answer received inconsistent 7/7 and 6/7 grades; grading limitation.

All evaluations and model links. Exact evaluated model revisions and SHA-256 manifests are included. Raw partial open-book traces remain local.