CoolFace
Modelpublic

dimitarpg13/semsimula-fock-parflm-anisogaussian-vtheta-owt-d768-gammasweep

sourceHugging Facecc-by-4.0updated 5h agoView on Hugging Face
0likes498downloads
Model Card

Fock-PARFLM v2.1 Anisotropic Gaussian V_theta + Fock Regularisation — Gamma Sweep with Geodesic Residual Analysis (OpenWebText, d=768)

## ⚠️ Correction — 2026-09-26 The residual as coded cannot detect a geodesic in this regime. Calibrated 2026-09-26 on a known geodesic — exact damped Newtonian dynamics in a smooth bounded potential, fed to this notebook's own conformal_grad / christoffel_vv — it reads \\(\bar{R}\\) between 0.84 and 1.06 at dt = 1 per layer with \\(\omega \Delta t \approx 1\\), which is the regime of this sweep. Two defects: the geodesic equation's reparametrisation term \\(\tfrac{\nabla V \cdot \dot{x}}{E - V} \dot{x}\\) is omitted, and \\(E\\) is frozen at layer 0. Even with both repaired, a second difference in layer index is not a derivative at one step per period, so no continuous-time residual can read near 0 here. Script: `geodesic_residual_calibration.py`. Measured here: \\(\bar{R}\\) ranges 1.281 to 3.579, with 0 of 8 checkpoints below the null. Every checkpoint is above the null. At no swept damping does adding the Christoffel and damping terms improve on simply ignoring them. These trajectories are not geodesics of the metric induced by their own potential. What still stands. The ranking results are unaffected, because they are claims about where the minimum falls, not about its absolute level: the PPL and \\(\bar{R}\\) minima do coincide at \\(\gamma = 0.050\\), and the damping-predictor comparisons hold. What does not stand is any reading of this sweep as demonstrating geodesic trajectories, or of the \\(\bar{R}\\)-minimising checkpoint as "geometrically faithful" in absolute terms. Lower is more geodesic-like than higher; that is a comparison, not a certificate. Two further cautions. These are 3,000-step checkpoints (PPL 327–393) — barely trained, so the geometry measured is close to that of the initialisation. And \\(\gamma{\text{geo}}\\) clusters near 0.96–0.98 independently of the training \\(\gamma\\) and of width. **This is an artefact of the diagnostic, not a property of the models:** on the known geodesic above, the same least-squares fit returns \\(\gamma{\text{geo}} = 0.93 / 0.82 / 0.75\\) for true damping \\(0.05 / 0.10 / 0.30\\) — the omitted reparametrisation term, which lies along the velocity, is what the fitted \\(\gamma\\) absorbs. The earlier "intrinsic preferred geometry" reading, and the "contraction rate" hedge in a previous revision of this correction, are both withdrawn. For contrast, a fully conservative Fock-PARFLM trained under the CfC+BAOAB integrator (no reverse channel, OpenWebText, d=384, L=2, 32,500 steps) has a layer step that is the damped \\(V\theta\\) geodesic step followed by LayerNorm, at a deflection of **0.0003** — three orders of magnitude below the null this sweep never reaches. See [`GeodesicExperimentswithCfCBAOAB.md`](https://github.com/dimitarpg13/semsimula-paper/blob/main/companionnotes/GeodesicExperimentswithCfCBAOAB.md) §4.9.

This repository holds eight short (3,000-step) training runs, one per candidate damping coefficient \\(\gamma \in \{0.05, 0.10, 0.15, 0.20, 0.25, 0.30, 0.40, 0.50\}\\), of the depth-conditioned anisotropic Gaussian \\(V_\theta\\) with Fock-coupling regularisation architecture, scaled up to d=768, L=16 and trained on OpenWebText. This is not a final trained model — it is the diagnostic sweep used to pick the damping coefficient for a subsequent full 100,000-step training run at this scale. Each of the 8 checkpoints is included in full.

Alongside the perplexity sweep, every checkpoint is also scored with the damped-geodesic residual \\(\bar{R}(\gamma)\\) — a closed-form diagnostic (no additional training, no autodiff through a learned metric) that measures how closely each model's own hidden-state trajectory follows a geodesic of the Riemannian (Jacobi) metric induced by its own learned potential. The headline result:

Both the perplexity minimum and the geodesic-residual minimum land at \\(\gamma = 0.050\\) — the smallest candidate tested, i.e. a boundary optimum rather than the interior minimum seen at d=384. This is also an exact, zero-parameter match to the two-regime closed-form damping predictor's high-\\(d\\) anchor prediction (\\(\gamma^\ast{\text{pred}} = 0.050\\)) — the first time this predictor has been confirmed on a bounded anisotropic-Gaussian \\(V\theta\\) at \\(d \geq 768\\), at a depth (\\(L=16\\)) the original MLP-\\(V_\theta\\) d=768 sweep never used. See Coincidence, Boundary, and an Exact Predictor Match.

Based on this sweep, a full 100,000-step run at \\(\gamma = 0.05\\) is expected to follow the same launch pattern as the d=384 run (currently in progress at \\(\gamma=0.10\\)); as of this card's creation the d=768 full run has not yet been launched.

This model is from the Semantic Simulation framework.

Table of Contents

When to Use This Repository

Use this repository if you want to:

  • —Reproduce or extend the gamma-selection methodology for the anisotropic-Gaussian + Fock-reg Fock-PARFLM line at d=768, L=16 on OpenWebText, including the geodesic-residual diagnostic.
  • —Study the non-monotonic PPL-vs-gamma wiggle that this width shares with the d=1024 sibling sweep but that no MLP-\\(V_\theta\\) sweep at comparable widths shows.
  • —*Compare relative geodesic fidelity across damping regimes* on a bounded, analytically-differentiable potential — the closed-form Jacobi-metric machinery here is structurally unavailable to attention-based or MLP-potential architectures.

Do not use this repository if you want a well-trained OpenWebText language model: every checkpoint here has seen only 3,000 steps (~24.6M tokens at effective batch 8 x block 512) and none is intended to produce fluent text. For a fully trained OpenWebText-scale Fock-PARFLM checkpoint, see semsimula-fock-parflm-depthcond-vtheta-openwebtext (27.23 PPL, isotropic Gaussian, d=384, 250K steps) — no full-length d=768 run in this line exists yet.

Architecture

Identical Fock-PARFLM v2.1 scaffold to the d=384 gamma-sweep sibling, scaled up to d=768 and, specific to this width, given explicit force bounding for numerical stability:

Input tokens x_1, ..., x_T
       |
   Untied token embedding E[x] + learned positional P[t]
       |
   For each of L=16 damped Störmer–Verlet integration steps (shared force field):
       |
       +-- K=5 causal-EMA context channels:
       |     xi^(m)_t = causal_ema(h, alpha_m)      [horizons ~2 .. ~200 tokens]
       |
       +-- Depth-conditioned multi-context V_theta (Anisotropic Gaussian, force-bounded):
       |     xi_g^(m) = xi^(m) + e_g^(m)                    [per-layer depth code]
       |     diff_k^m = h - mu_k^m(xi_g^(m))
       |     V_m = -sum_k w_k^m exp(-0.5[a_k^m . diff_k^{m2} + ||B_k^{mT} diff_k^m||^2])
       |     V_theta = sum_m V_m(xi_g^(m), h)               [5 contexts, 40 wells total]
       |     f_theta = clamp(-analytical_grad_h V_theta, max_norm=2/sqrt(768))
       |
       +-- Sparse pairwise V_phi (structural-competitive, 4 heads):
       |     top-k=16 past tokens per query (Gumbel routing)
       |     f_phi = -grad_h V_phi(h_t, h_s)                [autograd, sparse]
       |
       +-- Fock register pool (v2, 32 registers):
       |     M=32 virtual registers, Q/K/V creation gates, d_k=64
       |     LIFO stack discipline, per-register tau/keys, ortho init
       |     register repulsion (Gram penalty, lambda=0.05)
       |     reverse channel (per-layer, stabilised, pre-LN, soft-norm, warmup 4000)
       |     prefix-causal (leak-free by construction)
       |     f_fock = creation + destruction + exchange forces
       |
       +-- Total force: f = clamp(f_theta + f_phi + f_fock, max_norm=2/sqrt(768))
       |
       +-- Damped Stormer-Verlet step: h += (h-h_prev)/(1+dt*gamma) + dt^2*f/(m*(1+dt*gamma))
       |
       +-- LayerNorm(h)
       |
   Logits = h @ W_out^T + b_out                              [UNTIED W_out]

Auxiliary training-only loss term (not part of the forward pass above):
   L_fock_coupling = -lambda_fock * sum_k log(alpha_k + eps)   [log-barrier on xi coupling]
ParameterValue
Hidden dim (d)768
Layers (L)16
Max sequence length1024 (trained at block 512)
Vocab (GPT-2 BPE)50257
V_theta kindDepth-conditioned multi-context anisotropic Gaussian (bounded mixture)
Vtheta contexts (nctx)5 (one bank per xi channel)
Wells per context (K)8
Total attractors40
Anisotropic rank (r)4 — low-rank factor B_k in R^(768x4) per well
Precision init / capa_k init at -log(768) (log-precision), capped at 2/768
Force bounding (d768-specific)\\(V_\theta\\)'s own gradient norm and the model-level total force are both clamped to \\(2/\sqrt{768} \approx 0.0722\\) — not present in the d=384 sweep, added for stability at this width
Depth codesper-layer additive shift, shape (L=16, n_ctx=5, d=768), init std 0.02
Xi channels (K_xi)5 (alpha inits 0.50, 0.75, 0.95, 0.99, 0.995)
Fock-coupling regulariserlog-barrier on alpha_k, lambda=0.005, eps=1e-6 (see the TinyStories anchor for the full derivation)
V_phi kindstructural_competitive
Vphi heads / dtype / d_angle4 / 32 / 16
Top-k (sparse routing)16
Fock versionv2
Registers (M)32
Register d_k64
Stack disciplineLIFO
Reverse channelYes — per-layer, stabilised, pre-LN, soft-norm, 4,000-step warmup
Register repulsionGram penalty, lambda=0.05
EmbeddingsUntied (separate W_out)
Mass modellogfreq (frozen OpenWebText surprisal lookup)
lambdaV (Vtheta regularisation)0.01
Prefix-causal registersYes — every checkpoint trained natively leak-free from step 0
Total parameters225,354,011
V_theta parameters141,834,280 (63% of total — the low-rank correction \\(Bk \in \mathbb{R}^{d \times r}\\) scales with \\(d\\), so widening from d=384 to d=768 roughly quadruples \\(V\theta\\)'s share of the model)

The analytical form of \\(V\theta\\) — the diagonal-plus-rank-4 precision \\(\Sigmak^{-1} = \mathrm{diag}(ak) + Bk B_k^\top\\), the depth-conditioning mechanism, and the closed-form bounded gradient — is unchanged from the TinyStories anchor and the d=384 sibling; only \\(d\\) changes, plus the explicit force-norm clamp described above.

The Gamma Sweep

Eight candidate damping coefficients \\(\gamma \in \{0.05, 0.10, 0.15, 0.20, 0.25, 0.30, 0.40, 0.50\}\\) were each trained from scratch for 3,000 steps (WSD schedule, peak LR 3e-4, effective batch 8, up to 1B-token training pool), then scored on a held-out 2M-token OpenWebText validation slice. This protocol mirrors the d=384 sweep and is designed to be cheap: 8 x 3,000 steps rather than 8 x 100,000. The batch/grad-accum split (2x4 here vs. 2x8 at d=384) and the tighter gradient-clip ceilings (0.5 global / 0.2 for V_phi, vs. 1.0 / 0.3 at d=384) are both d768-specific stability adjustments.

The Geodesic Residual Diagnostic

Fock-PARFLM's scalar potential \\(V\theta\\) is a closed-form Gaussian mixture with an analytical gradient, which makes a diagnostic available here that is structurally unavailable to attention-based or MLP-potential architectures: at fixed energy \\(E\\), Hamiltonian trajectories are geodesics of the **Jacobi metric** \\(g^J{ij}(x) = 2(E - V(x))\delta{ij}\\), a conformally flat metric whose Christoffel symbols are closed-form functions of \\(\nabla V\theta\\) — no learned metric, no autodiff through a metric, no boundary-value solve.

For a trajectory with position stream \\(x\ell\\), velocity stream \\(v\ell\\), and measured acceleration \\(a\ell\\) (the discrete second difference of \\(x\ell\\), consistent with the model's Störmer–Verlet integrator, whose velocity is the position difference), the per-layer damped-geodesic residual is

$$ R\ell = \frac{\big\lVert a\ell + \Gamma(v\ell, v\ell) + \gamma v\ell \big\rVert}{\lVert a\ell \rVert + \varepsilon}, \qquad \Gamma(v,v)^k = \Gamma^k_{ij} v^i v^j, $$

where \\(\Gamma\\) is computed in closed form from \\(V\theta\\)'s analytical gradient. \\(R\ell \approx 0\\) means the trajectory is a damped geodesic of the metric induced by the model's own learned potential — this is not a pure-conservation claim (the explicit \\(\gamma v\ell\\) damping term is included), only that the dynamics satisfy the damped geodesic equation with the architecture's own damping coefficient. Averaging over layers and 10 fixed validation batches (seed 42) gives \\(\bar{R}(\gamma{\text{train}})\\), evaluated at \\(\gamma{\text{eval}} = \gamma{\text{train}}\\) for each retained checkpoint — the diagonal overlay against \\(\mathrm{PPL}(\gamma_{\text{train}})\\).

How to read the scale — added 2026-09-26. The residual is normalised by the bare acceleration \\(\lVert a_\ell \rVert\\), which fixes three reference points:

\\(\bar{R}\\)meaning
\\(\approx 0\\)geodesic — the geometry accounts for the acceleration
\\(\approx 1\\)the null. \\(\Gamma(v,v) + \gamma v\\) contributes nothing: the residual is the same size as the acceleration you started with
\\(> 1\\)the geometric terms make the fit worse than omitting them

This scale was not stated in the original version of this card, and it changes how the table below should be read. See Correction (2026-09-26) near the top.

A second, closed-form quantity — the recovered intrinsic damping \\(\gamma_{\text{geo}}\\) — is the least-squares row minimiser "the damping value that best explains this specific trajectory," independent of what \\(\gamma\\) the model was actually trained with:

$$ \gamma{\text{geo}} = -\frac{\big\langle a\ell + \Gamma(v\ell, v\ell),\ v\ell \big\rangle}{\lVert v\ell \rVert^{2}}. $$

Full derivation, practical mitigations (turning-point exclusion, reference-energy convention, integrator staggering), and validation controls (vanilla-baseline, shuffled-\\(\Gamma\\), and random-direction nulls) are in the companion note `Geodesic_Preservation_Experiment.md`.

Results: A Boundary Optimum at Gamma 0.05

\\(\gamma\\)PPL\\(\bar{R}\\)\\(\gamma_{\text{geo}}\\)Excluded frac
0.050326.97 ← PPL min1.2813 ← \\(\bar{R}\\) min0.98290%
0.100350.701.96480.98280%
0.150340.251.51590.98270%
0.200368.621.75440.97950%
0.250362.392.90920.98400%
0.300362.441.49010.97350%
0.400393.473.57900.98480%
0.500364.942.67500.97810%

<p align="center"><img src="results/geodesicoverlayanisogaussiand768.png" alt="PPL vs geodesic residual overlay for d=768 aniso-Gaussian + fock-reg gamma sweep" width="660"></p> The minima coincide, at the boundary of the swept range. Both PPL and \\(\bar{R}\\) bottom out at \\(\gamma = 0.050\\), the smallest gamma tested — unlike the d=384 sweep, where the joint minimum sat at an interior point (\\(0.100\\)) with worse candidates on both sides. \\(\gamma_{\text{geo}}\\) also shows the tightest cross-gamma clustering of any sweep in this family to date: mean \\(0.9810\\), std \\(0.0035\\) — essentially flat across a full order-of-magnitude range of nominal training \\(\gamma\\) (0.05 to 0.50), the same "intrinsic preferred geometry" signature documented at every other scale, but here with less than half the spread seen at d=384.

The per-layer residual heatmap below shows where in the network the departures from geodesic behaviour concentrate at each damping level — notice the bright band across nearly every layer at \\(\gamma=0.40\\), the sweep's worst PPL candidate, and the uniformly dark (near-geodesic) row at \\(\gamma=0.05\\):

<p align="center"><img src="results/geodesicperlayeranisogaussian_d768.png" alt="Per-layer geodesic residual heatmap for d=768 aniso-Gaussian + fock-reg gamma sweep" width="660"></p> The recovered intrinsic damping is visually almost a flat line across the entire swept range:

<p align="center"><img src="results/gammageorecoveryanisogaussian_d768.png" alt="Recovered intrinsic damping vs training gamma for d=768 aniso-Gaussian + fock-reg gamma sweep" width="660"></p>

The Non-Monotonic Wiggle Past the Minimum

Unlike the original MLP-\\(V\theta\\) d=768 sweep (a clean, monotonic PPL increase from \\(\gamma=0.05\\) to \\(0.50\\); see [`Fock-PARFLMScale-UpGammaSweepResultsandDampingRegimeAnalysis.md`](https://github.com/dimitarpg13/semsimula-paper/blob/main/companionnotes/Fock-PARFLMScale-UpGammaSweepResultsandDampingRegimeAnalysis.md)), this anisotropic-Gaussian sweep is non-monotonic past its minimum: PPL rises from 326.97 (\\(\gamma=0.05\\)) to a local peak at \\(\gamma=0.20\\) (368.62), dips back down at \\(\gamma=0.25\\)-\\(0.30\\) (362.39, 362.44), rises again to the sweep's worst point at \\(\gamma=0.40\\) (393.47), then partially recovers at \\(\gamma=0.50\\) (364.94).

The overall ranking is not in doubt — \\(\gamma=0.05\\) leads its nearest competitor (\\(\gamma=0.15\\), 340.25 PPL) by 13.28 PPL, about 4% — but the wiggle itself is a genuine feature of the aniso-Gaussian + Fock-reg configuration that no MLP-\\(V\theta\\) sweep at comparable widths shows. A companion sweep at d=1024 (not yet uploaded to this collection) shows the **same wiggle shape**, with a local peak in the same \\(\gamma \approx 0.20\)-\\(0.40\\) range before a partial recovery at 0.40-0.50 — a weak but real argument against pure single-run noise, since independent per-candidate noise would not obviously line up at matching \\(\gamma\\) values across two different widths. Two candidate explanations remain open (not mutually exclusive): (1) single-seed optimisation noise riding on a genuinely flat or slowly-varying underlying PPL surface, or (2) a structural interaction between the anisotropic precision matrices (or the reverse channel) and mid-range explicit friction. Disentangling them would require a 2-3 seed rerun of one or two candidates (e.g. \\(\gamma=0.30\\)) — not yet done. See [`DeterminingoptimalgammaforFock-PARFLM.md` §14.5](https://github.com/dimitarpg13/semsimula-paper/blob/main/companionnotes/DeterminingoptimalgammaforFock-PARFLM.md#145-cross-scale-comparison-a-shared-non-monotonic-signature-at-d768) for the full cross-scale discussion.

Coincidence, Boundary, and an Exact Predictor Match

The two-regime closed-form damping predictor, calibrated on the original (MLP) \\(V\theta\\) family, gives at \\(d=768\\), \\(L=16\\), \\(\bar{m}=1.4\\) (using the high-\\(d\\) regime constant \\(\rho{\text{hi}} = 0.565\\)):

$$ \gamma^{\ast}_{\text{pred}} = \frac{1.4}{16}\ln(1/0.565) = 0.0875 \times 0.571 = 0.050 $$

Empirical \\(\gamma^{\ast} = 0.050\\). Exact match, zero error. This resolves an open question left by the d=384 sweep — whether the aniso-Gaussian family's crossover, once past \\(d=384\\), matches the high-\\(d\\) anchor exactly (as the MLP-\\(V\theta\\) family does) or retains a residual offset. It matches exactly, and at a depth (\\(L=16\\)) the original MLP \\(d=768\\) sweep never tested (that sweep used \\(L=12\\), where the depth-only formula overshoots by 33% — a known, pre-existing discrepancy unrelated to this architecture). Combined with the same exact match found in the d=1024 companion sweep, this is the strongest evidence to date that \\(\rho{\text{hi}}=0.565\\) is a genuine architecture-family invariant at \\(d \gtrsim 768\\), not an artefact of the MLP \\(V_\theta\\) or of the specific depths swept when the constant was first calibrated.

Sweep\\(d\\)\\(L\\)\\(V_\theta\\)\\(\gamma^\ast_{\text{pred}}\\)\\(\gamma^\ast_{\text{empirical}}\\)Match
Original MLP d=768 sweep76812MLP0.0670.05+33% (known \\(L=12\\) depth-formula overshoot)
This sweep76816aniso-Gaussian + fock-reg0.0500.050Exact
d=1024 companion (not yet uploaded)102416aniso-Gaussian + fock-reg0.0500.050Exact

Comparison to the d=384 Sibling

Sweep\\(\gamma^\ast_{\text{PPL}}\\)\\(\gamma^\ast_{\bar{R}}\\)Coincidence?Optimum shapeMargin vs. runner-up
d=384, L=160.1000.100Yes — gap 0Interior minimumlarge
d=768, L=16 (this sweep)0.0500.050Yes — gap 0Boundary minimum~4%, modest

Both widths show coinciding PPL and geodesic-residual minima — extending the pattern first seen at d=384 to a second, larger width — but the shape of the optimum differs: d=384's winner sits strictly between two worse neighbours (a genuine interior bowl, only slightly undermined by the \\(\gamma=0.20\\) divergence outlier there), while d=768's winner is the smallest gamma tested, with PPL and \\(\bar{R}\\) both rising (non-monotonically) as gamma increases from there. The \\(\gamma_{\text{geo}}\\) clustering is also tighter at d=768 (std 0.0035 vs. a wider spread at d=384, driven partly by that sweep's \\(\gamma=0.20\\) outlier) — consistent with a model that has, if anything, an even more uniform intrinsic damping preference at this width.

Caveats: Short-Sweep Reliability

This is a 3,000-step, single-seed measurement per gamma, and the family has one documented case where a short-sweep ranking reversed at full training length: the d=256 TinyStories aniso-Gaussian sweep favoured \\(\gamma=0.150\\) at 3K steps, but the full 20K-step run favoured \\(\gamma=0.300\\) instead. Two considerations specific to this sweep are worth weighing before trusting \\(\gamma=0.05\\) at a full 100K-step horizon:

  • —The margin between \\(\gamma=0.05\\) and its nearest competitor (\\(\gamma=0.15\\)) is real but modest — about 4%, right at the edge of the family's informal "flat bowl, ranking unreliable" threshold (~5%) from the earlier d=256 reversal analysis. The d=1024 companion sweep's margin (~10-11%) is comfortably wider.
  • —The non-monotonic wiggle documented above shows this configuration is not perfectly smooth in gamma, so the training log for a full run should still be watched for instability signatures analogous to the d=384 sweep's \\(\gamma=0.20\\) divergence-and-partial-recovery excursion, as an early-warning check.

Two considerations nonetheless favour trusting this sweep's recommendation more than a typical modest-margin case: (a) unlike the flat interior bowls at \\(d \leq 384\\) where the documented reversal actually occurred, this is a boundary optimum with \\(\gamma_{\text{geo}}\\) essentially flat across the whole range — there is no competing interior basin the sweep could be misranking within, only a preference for less explicit friction that a longer horizon is unlikely to overturn; and (b) the predictor's independent, zero-parameter prediction lands on the exact same value (above). Recommendation: \\(\gamma=0.05\\) for the d=768, L=16 aniso-Gaussian full run.

How to Get Started

python
import math, torch, sys
sys.path.insert(0, "multixi")
sys.path.insert(0, "parf")
sys.path.insert(0, "energetic_minima")
sys.path.insert(0, "sarf_mass_variant")

from parf.model_fock_parf_multixi import FockMultiXiPARFLM, FockMultiXiPARFConfig
from parf.model_aniso_gaussian_vtheta import AnisotropicDepthConditionedGaussianVTheta, install_aniso_depth_routing
from huggingface_hub import hf_hub_download

REPO = "dimitarpg13/semsimula-fock-parflm-anisogaussian-vtheta-owt-d768-gammasweep"
GAMMA = "0.050"  # the recommended candidate; also available: 0.100, 0.150, 0.200, 0.250, 0.300, 0.400, 0.500

logfreq_path = hf_hub_download(repo_id=REPO, filename="results/logfreq_surprisal_openwebtext.npy")
FORCE_MAX = 2.0 / math.sqrt(768)

config = FockMultiXiPARFConfig(
    vocab_size=50257, d=768, max_len=1024, L=16,
    v_hidden=1024, v_depth=3, dt=1.0,
    mass_mode="logfreq", logfreq_path=logfreq_path, logfreq_init_alpha=0.1,
    init_gamma=1.0, fixed_gamma=float(GAMMA),
    causal_force=True, ln_after_step=True,
    xi_channels=5, xi_alpha_inits=[0.50, 0.75, 0.95, 0.99, 0.995],
    xi_learnable=True, xi_alpha_init_mode="explicit",
    fock_version="v2", n_registers=32,
    reverse_channel=True, reverse_channel_stable=True, reverse_channel_pre_ln=True,
    reverse_channel_soft_norm=True, reverse_channel_warmup_steps=4000, reverse_channel_per_layer=True,
    register_repulsion=True, register_repulsion_coeff=0.05,
    prefix_causal_registers=True,
    v_phi_kind="structural_competitive", v_phi_n_heads=4, v_phi_d_type=32, v_phi_d_angle=16,
    v_phi_eps=0.1, v_phi_phi_hidden=128, v_phi_theta_hidden=128, v_phi_mlp_hidden=128,
    top_k=16,
    use_output_bias=True, tie_embeddings=False,
    score_head_hidden=32,
    gumbel_tau_init=1.0, gumbel_tau_min=0.3, gumbel_noise=True,
    use_gathered_v_phi=True,
    use_layer_checkpoint=True,
    ln_before_distance=True, per_layer_v_phi_scale=True,
    register_salience_decay=0.5, register_salience_threshold=0.005,
    creation_gate_hidden=64, stack_discipline=True,
    d_k=64, tau_create_init=8.0,
    per_register_tau=True, per_register_keys=True,
    ortho_register_init=True,
    force_clamp_max=FORCE_MAX,
)
model = FockMultiXiPARFLM(config)

model.V_theta = AnisotropicDepthConditionedGaussianVTheta(
    d=768, K=8, n_ctx=5, n_layers=16, rank=4,
    w_scale=1.0,
    init_log_precision=-math.log(768),
    precision_max=2.0 / 768,
    force_norm_max=FORCE_MAX,
    code_init_std=0.02,
)
install_aniso_depth_routing(model)

ckpt_path = hf_hub_download(repo_id=REPO, filename=f"checkpoints/gamma_{GAMMA}/ckpt_best.pt")
state = torch.load(ckpt_path, map_location="cpu", weights_only=False)
model.load_state_dict(state["model_state_dict"])
model.eval()

print(f"Parameters: {sum(p.numel() for p in model.parameters()):,}")            # 225,354,011
print(f"gamma={state['gamma']}  step={state['step']:,}  val_ppl={state['val_ppl']:.2f}")

Available Artifacts

PathDescription
checkpoints/gamma_0.050/ckpt_best.pt ... checkpoints/gamma_0.500/ckpt_best.ptAll 8 gamma-sweep checkpoints (step 3,000 each), each a dict with model_state_dict, optimizer_state_dict, step, val_loss, val_ppl, gamma, v_theta_variant, aniso_rank, lambda_fock_reg
checkpoints/gamma_*/training_log.jsonlPer-500-step training metrics and the step-2,000 causal_probe event, one file per gamma
sweep_summary.jsonSweep-level summary (best PPL per gamma, wall-clock)
results/geodesic_results.jsonFull per-gamma geodesic residual data: \\(\bar{R}\\), \\(\gamma_{\text{geo}}\\), excluded fraction, and all 15 per-layer residuals
results/geodesic_overlay_aniso_gaussian_d768.pngPPL vs. geodesic residual dual-axis overlay
results/geodesic_per_layer_aniso_gaussian_d768.pngPer-layer, per-gamma geodesic residual heatmap
results/gamma_geo_recovery_aniso_gaussian_d768.pngRecovered intrinsic damping \\(\gamma_{\text{geo}}\\) vs. training gamma
results/logfreq_surprisal_openwebtext.npyFrozen per-token log-frequency (surprisal) lookup used by the logfreq mass model
model_aniso_gaussian_vtheta.pyAnisotropic Gaussian Vtheta classes + `installanisodepthrouting` (identical file to the TinyStories anchor's and the d=384 sibling's)
config.jsonFull sweep configuration, per-gamma results table, and predictor comparison

Training Details

Training Data

OpenWebText, tokenized with GPT-2 BPE (vocab 50257). Each candidate trains on up to 1B tokens (early stopped at 3,000 steps, effective batch 8, block 512 -> ~12.3M tokens actually consumed) and is evaluated on a held-out 2M-token validation slice with no train/val overlap.

Training Procedure (per gamma candidate)

HyperparameterValue
OptimizerAdamW (betas 0.9/0.95, weight decay 0.01)
LR scheduleWSD (warmup 0-150, stable 150-1950, decay 1950-3000)
Peak learning rate3e-4
Batch size2 x grad-accum 4 (effective 8)
Block size512
Steps3,000
Gradient clippingper-group (global 0.5; V_phi 0.2 — tighter than d=384's 1.0/0.3, a d768-specific stability adjustment)
Force clamp (model-level and V_theta-level)\\(2/\sqrt{768} \approx 0.0722\\)
lambdaV (Vtheta regularisation)0.01
lambda_fock (coupling regulariser)0.005
Seed0
Hardware1x NVIDIA H100/A100 (Google Colab)

Causal-Leak Verification

All 8 checkpoints were trained natively with prefix_causal_registers=True from step 0. The bit-exact future-perturbation causal probe passed with max_delta=0.0 at step 2,000 for every one of the 8 gamma candidates.

Training Script

notebooks/conservative_arch/scaleup/colab_fock_gamma_sweep_geodesic_aniso_gaussian_fockreg_d768.ipynb (companion repo) — self-contained Colab notebook that runs the 8-candidate sweep, the geodesic residual analysis, and produces the overlay/heatmap/recovery figures in one pass, requiring no additional training beyond the sweep itself.

Evaluation Results

OpenWebText Validation Perplexity (3,000-step sweep candidates)

\\(\gamma\\)PPLRank
0.050326.971st
0.150340.252nd
0.300362.443rd
0.250362.394th
0.500364.945th
0.100350.706th
0.200368.627th
0.400393.478th
These PPL values are from 3,000-step short-sweep candidates and are not comparable to the fully trained OpenWebText checkpoints elsewhere in this family (e.g. 27.23 PPL after 250K steps for the isotropic d=384 flagship). They exist solely to rank candidate damping coefficients. Note the ranking is not simply increasing with \\(\gamma\\); see The Non-Monotonic Wiggle Past the Minimum.

See Results: A Boundary Optimum at Gamma 0.05 for the combined PPL / geodesic-residual table.

SPLM Family Overview

This model is part of the Semantic Simulation SPLM family:

ModelDesignCorpusPPLHuggingFace
Multi-Xi SPLM (MLP)Pure scalar potentialTinyStories11.51semsimula-splm-multixi
Multi-Xi PARFLM (MLP)Scalar + pairwise forcesTinyStories12.06semsimula-parflm-multixi
Fock-PARFLM v2.1 (MLP)PARFLM + Fock registersTinyStories9.70semsimula-fock-parflm
Fock-PARFLM v2.1 (SQ3)Structured + pairwise + FockTinyStories10.90semsimula-fock-parflm-structured-vtheta
Fock-PARFLM v2.1 (depth-cond. isotropic Gaussian)Bounded multi-context + pairwise + FockTinyStories16.33semsimula-fock-parflm-depthcond-vtheta
Fock-PARFLM v2.1 (depth-cond. anisotropic Gaussian + fock-reg)Bounded, ellipsoidal multi-context + pairwise + FockTinyStories9.04semsimula-fock-parflm-anisogaussian-vtheta
Fock-PARFLM v2.1 (aniso-Gaussian + fock-reg, first-order/Fock-G1)Same as above, gradient-flow integratorTinyStories8.95semsimula-fock-parflm-anisogaussian-vtheta-fock-g1
Fock-PARFLM v2.1 (depth-cond. Gaussian, OpenWebText scale)Bounded multi-context + pairwise + Fock, d=384 L=16OpenWebText27.23semsimula-fock-parflm-depthcond-vtheta-openwebtext
Fock-PARFLM v2.1 (aniso-Gaussian + fock-reg, gamma sweep, d=384)Bounded, ellipsoidal multi-context + pairwise + Fock — 8-way gamma sweep + geodesic analysisOpenWebText278.27 (best of 8, 3K-step sweep, not a final model)semsimula-fock-parflm-anisogaussian-vtheta-owt-d384-gammasweep
Fock-PARFLM v2.1 (aniso-Gaussian + fock-reg, gamma sweep, d=768)Bounded, ellipsoidal multi-context + pairwise + Fock — 8-way gamma sweep + geodesic analysisOpenWebText326.97 (best of 8, 3K-step sweep, not a final model)this repository
Fock-AttentionFock + attentionTinyStories9.42semsimula-fock-attention
Hybrid SPLM+AttnAttention + SPLM refinementTinyStories8.50semsimula-hybrid-splm

Collection: Semantic Simulation SPLM Model Family

Bias, Risks, and Limitations

  • —Not a final trained model. Every checkpoint in this repository has been trained for only 3,000 steps — a gamma-selection diagnostic, not a language model intended for generation or downstream use. Do not compare its PPL to fully trained checkpoints elsewhere in this family.
  • —Single-seed sweep. Each gamma candidate is one run; no seed-variance estimate is available. The margin between \\(\gamma=0.05\\) and its nearest competitor (~4%) is real but modest; see Caveats: Short-Sweep Reliability.
  • —Short-horizon ranking can reverse. The family has one documented case (d=256 TinyStories) where the short-sweep-optimal gamma differed from the full-run-optimal gamma. The \\(\gamma=0.05\\) recommendation here carries some of the same risk, mitigated by the boundary-optimum shape and the exact predictor match; see the caveats section.
  • —Non-monotonic PPL-vs-gamma wiggle, shared with the d=1024 companion sweep, is not yet explained (single-seed noise vs. a structural property of this \\(V_\theta\\) configuration); see above.
  • —OpenWebText only, English only. No instruction tuning, no RLHF/DPO, no safety filtering.
  • —Geodesic residual is a diagnostic, not a training objective. The model was trained to minimise cross-entropy; the near-geodesic behaviour at \\(\gamma=0.05\\) is an emergent structural property, not something the loss function directly optimises for.
  • —Very large \\(V_\theta\\) hypernetwork. At 141.8M parameters, \\(V\theta\\) alone is 63% of the 225M-parameter total — a substantially larger share than the d=384 sibling's 46%, driven by the low-rank correction \\(Bk \in \mathbb{R}^{d \times r}\\) scaling with \\(d\\).
  • —No causal-leak issue. All 8 checkpoints were trained natively with prefix_causal_registers=True; see Causal-Leak Verification.

Citation

bibtex
@misc{Gueorguiev2026SemSim,
  author    = {Gueorguiev, Dimitar P.},
  title     = {Semantic Simulation: A Prescriptive Lagrangian Framework
               for Efficient Semantic Inference --- A Conservative-by-
               Construction Language Model and the Shared-Potential
               Separator, with a Correspondence to Joint Embedding
               Predictive Architectures},
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.19712427},
  url       = {https://doi.org/10.5281/zenodo.19712427},
  note      = {Companion code repository:
               \url{https://github.com/dimitarpg13/semsimula-paper}}
}

Environmental Impact

  • —Hardware: 1x NVIDIA H100/A100 80GB (Google Colab)
  • —Training: 8 x 3,000 steps = 24,000 total training steps across the sweep, plus inference-only geodesic residual analysis (10 validation batches x 8 checkpoints, no additional training)
  • —Carbon footprint: small; a single-GPU research sweep, estimated on the order of a few kg CO2
Correction (2026-09-27) — integrator name. This model's layer step was described here as a damped Euler step. It is not. The update is h_new = h + (h - h_prev)/(1 + dt*gamma) + dt^2*f/(m*(1 + dt*gamma)), which carries no velocity state at all: the velocity is the position difference h - h_prev. Undamped this is h_{n+1} = 2*h_n - h_{n-1} + dt^2*f/m, i.e. Störmer–Verlet in position form. It is also not velocity-Verlet, which carries an explicit velocity through half-kick/drift/half-kick. The distinction matters because the \\(\omega \cdot dt < 2\\) stability wall is the Störmer/leapfrog bound and applies to this lineage; the semi-implicit Euler models in the same collection were never subject to it. Only the name was wrong — no measurement on this card changes.