dimitarpg13/semsimula-fock-parflm-anisogaussian-vtheta-owt-d384-gammasweep
Fock-PARFLM v2.1 Anisotropic Gaussian V_theta + Fock Regularisation — Gamma Sweep with Geodesic Residual Analysis (OpenWebText, d=384)
## ⚠️ Correction — 2026-09-26 The residual as coded cannot detect a geodesic in this regime. Calibrated 2026-09-26 on a known geodesic — exact damped Newtonian dynamics in a smooth bounded potential, fed to this notebook's ownconformal_grad/christoffel_vv— it reads \\(\bar{R}\\) between 0.84 and 1.06 at dt = 1 per layer with \\(\omega \Delta t \approx 1\\), which is the regime of this sweep. Two defects: the geodesic equation's reparametrisation term \\(\tfrac{\nabla V \cdot \dot{x}}{E - V} \dot{x}\\) is omitted, and \\(E\\) is frozen at layer 0. Even with both repaired, a second difference in layer index is not a derivative at one step per period, so no continuous-time residual can read near 0 here. Script: `geodesic_residual_calibration.py`. Measured here: \\(\bar{R}\\) ranges 0.671 to 1.540, with 4 of 8 checkpoints below the null. Mixed. Four of the eight checkpoints sit below the null and four above. The best, 0.671 at gamma=0.100, means the geometry accounts for roughly a third of the measured acceleration -- a real but partial fit, not a geodesic. What still stands. The ranking results are unaffected, because they are claims about where the minimum falls, not about its absolute level: the PPL and \\(\bar{R}\\) minima do coincide at \\(\gamma = 0.100\\), and the damping-predictor comparisons hold. What does not stand is any reading of this sweep as demonstrating geodesic trajectories, or of the \\(\bar{R}\\)-minimising checkpoint as "geometrically faithful" in absolute terms. Lower is more geodesic-like than higher; that is a comparison, not a certificate. Two further cautions. These are 3,000-step checkpoints (PPL 244–632) — barely trained, so the geometry measured is close to that of the initialisation. And \\(\gamma{\text{geo}}\\) clusters near 0.96–0.98 independently of the training \\(\gamma\\) and of width. **This is an artefact of the diagnostic, not a property of the models:** on the known geodesic above, the same least-squares fit returns \\(\gamma{\text{geo}} = 0.93 / 0.82 / 0.75\\) for true damping \\(0.05 / 0.10 / 0.30\\) — the omitted reparametrisation term, which lies along the velocity, is what the fitted \\(\gamma\\) absorbs. The earlier "intrinsic preferred geometry" reading, and the "contraction rate" hedge in a previous revision of this correction, are both withdrawn. For contrast, a fully conservative Fock-PARFLM trained under the CfC+BAOAB integrator (no reverse channel, OpenWebText, d=384, L=2, 32,500 steps) has a layer step that is the damped \\(V\theta\\) geodesic step followed by LayerNorm, at a deflection of **0.0003** — three orders of magnitude below the null this sweep never reaches. See [`GeodesicExperimentswithCfCBAOAB.md`](https://github.com/dimitarpg13/semsimula-paper/blob/main/companionnotes/GeodesicExperimentswithCfCBAOAB.md) §4.9.
This repository holds eight short (3,000-step) training runs, one per candidate damping coefficient \\(\gamma \in \{0.05, 0.10, 0.15, 0.20, 0.25, 0.30, 0.40, 0.50\}\\), of the depth-conditioned anisotropic Gaussian \\(V_\theta\\) with Fock-coupling regularisation architecture, scaled up to d=384, L=16 and trained on OpenWebText. This is not a final trained model — it is the diagnostic sweep used to pick the damping coefficient for the subsequent full 100,000-step training run at this scale. Each of the 8 checkpoints is included in full.
Alongside the perplexity sweep, every checkpoint is also scored with the damped-geodesic residual \\(\bar{R}(\gamma)\\) — a closed-form diagnostic (no additional training, no autodiff through a learned metric) that measures how closely each model's own hidden-state trajectory follows a geodesic of the Riemannian (Jacobi) metric induced by its own learned potential. The headline result:
Both the perplexity minimum and the geodesic-residual minimum land at the same value, \\(\gamma = 0.100\\). The damping coefficient that produces the best language model is, on this architecture and corpus, also the one whose trajectory is least far from geodesic (see the Correction above: least far is not the same as faithful) — the same qualitative finding previously confirmed at d=768 and d=1024 with this \\(V\theta\\) family, but here demonstrated for the first time at d=384, where the isotropic/original-\\(V\theta\\) sweep on this exact width previously showed the opposite result (minima 5x apart; see Coincidence or Crossover?).
Based on this sweep, the full 100,000-step run (colab_fock_aniso_gaussian_fockreg_openwebtext.ipynb) was launched with \\(\gamma = 0.10\\); as of this card's creation that run is in progress and will be published as its own model card once complete.
This model is from the Semantic Simulation framework.
Table of Contents
- When to Use This Repository
- Architecture
- The Gamma Sweep
- The Geodesic Residual Diagnostic
- Results: Minima Coincide at Gamma 0.10
- The Gamma=0.20 Anomaly
- Coincidence or Crossover? Comparing Against the Isotropic Sibling at the Same Width
- Caveats: Short-Sweep Reliability
- How to Get Started
- Available Artifacts
- Training Details
- Evaluation Results
- SPLM Family Overview
- Bias, Risks, and Limitations
- Citation
- Environmental Impact
When to Use This Repository
Use this repository if you want to:
- Reproduce or extend the gamma-selection methodology for the anisotropic-Gaussian + Fock-reg Fock-PARFLM line at d=384, L=16 on OpenWebText, including the geodesic-residual diagnostic.
- Study any individual gamma candidate's short-horizon behaviour, e.g. the gamma=0.20 divergence-and-partial-recovery excursion (see below).
- *Compare relative geodesic fidelity across damping regimes* on a bounded, analytically-differentiable potential — the closed-form Jacobi-metric machinery here is structurally unavailable to attention-based or MLP-potential architectures.
Do not use this repository if you want a well-trained OpenWebText language model: every checkpoint here has seen only 3,000 steps (~50M tokens at effective batch 16 x block 512) and none is intended to produce fluent text. For a fully trained OpenWebText-scale Fock-PARFLM checkpoint, see semsimula-fock-parflm-depthcond-vtheta-openwebtext (27.23 PPL, isotropic Gaussian, 250K steps) — the full 100K-step run using the gamma selected here will supersede it for the anisotropic line once complete.
Architecture
Identical Fock-PARFLM v2.1 scaffold to the TinyStories anisotropic-Gaussian anchor, scaled up to OpenWebText width/depth and widened to 5 xi-context channels (matching the isotropic OpenWebText flagship's "5long" configuration):
Input tokens x_1, ..., x_T
|
Untied token embedding E[x] + learned positional P[t]
|
For each of L=16 damped Störmer–Verlet integration steps (shared force field):
|
+-- K=5 causal-EMA context channels:
| xi^(m)_t = causal_ema(h, alpha_m) [horizons ~2 .. ~200 tokens]
|
+-- Depth-conditioned multi-context V_theta (Anisotropic Gaussian):
| xi_g^(m) = xi^(m) + e_g^(m) [per-layer depth code]
| diff_k^m = h - mu_k^m(xi_g^(m))
| V_m = -sum_k w_k^m exp(-0.5[a_k^m . diff_k^{m2} + ||B_k^{mT} diff_k^m||^2])
| V_theta = sum_m V_m(xi_g^(m), h) [5 contexts, 40 wells total]
| f_theta = -analytical_grad_h V_theta [closed-form, bounded]
|
+-- Sparse pairwise V_phi (structural-competitive, 4 heads):
| top-k=16 past tokens per query (Gumbel routing)
| f_phi = -grad_h V_phi(h_t, h_s) [autograd, sparse]
|
+-- Fock register pool (v2, 32 registers):
| M=32 virtual registers, Q/K/V creation gates, d_k=64
| LIFO stack discipline, per-register tau/keys, ortho init
| register repulsion (Gram penalty, lambda=0.05)
| reverse channel (per-layer, stabilised, pre-LN, soft-norm, warmup 4000)
| prefix-causal (leak-free by construction)
| f_fock = creation + destruction + exchange forces
|
+-- Total force: f = f_theta + f_phi + f_fock
|
+-- Damped Stormer-Verlet step: h += (h-h_prev)/(1+dt*gamma) + dt^2*f/(m*(1+dt*gamma))
|
+-- LayerNorm(h)
|
Logits = h @ W_out^T + b_out [UNTIED W_out]
Auxiliary training-only loss term (not part of the forward pass above):
L_fock_coupling = -lambda_fock * sum_k log(alpha_k + eps) [log-barrier on xi coupling]The analytical form of \\(V\theta\\) — the diagonal-plus-rank-4 precision \\(\Sigmak^{-1} = \mathrm{diag}(ak) + Bk B_k^\top\\), the depth-conditioning mechanism, and the closed-form bounded gradient — is unchanged from the TinyStories anchor; only \\(d\\), \\(L\\), and the number of contexts (5 instead of 4) change.
The Gamma Sweep
Eight candidate damping coefficients \\(\gamma \in \{0.05, 0.10, 0.15, 0.20, 0.25, 0.30, 0.40, 0.50\}\\) were each trained from scratch for 3,000 steps (WSD schedule, peak LR 3e-4, effective batch 16, ~1B-token training pool), then scored on a held-out 2M-token OpenWebText validation slice. This protocol mirrors the gamma-sweep-then-full-run methodology used throughout the family (see e.g. the d=768 and d=1024 aniso-Gaussian sweeps) and is designed to be cheap: 8 x 3,000 steps rather than 8 x 100,000.
The Geodesic Residual Diagnostic
Fock-PARFLM's scalar potential \\(V\theta\\) is a closed-form Gaussian mixture with an analytical gradient, which makes a diagnostic available here that is structurally unavailable to attention-based or MLP-potential architectures: at fixed energy \\(E\\), Hamiltonian trajectories are geodesics of the **Jacobi metric** \\(g^J{ij}(x) = 2(E - V(x))\delta{ij}\\), a conformally flat metric whose Christoffel symbols are closed-form functions of \\(\nabla V\theta\\) — no learned metric, no autodiff through a metric, no boundary-value solve.
For a trajectory with position stream \\(x\ell\\), velocity stream \\(v\ell\\), and measured acceleration \\(a\ell\\) (the discrete second difference of \\(x\ell\\), consistent with the model's Störmer–Verlet integrator, whose velocity is the position difference), the per-layer damped-geodesic residual is
$$ R\ell = \frac{\big\lVert a\ell + \Gamma(v\ell, v\ell) + \gamma v\ell \big\rVert}{\lVert a\ell \rVert + \varepsilon}, \qquad \Gamma(v,v)^k = \Gamma^k_{ij} v^i v^j, $$
where \\(\Gamma\\) is computed in closed form from \\(V\theta\\)'s analytical gradient. \\(R\ell \approx 0\\) means the trajectory is a damped geodesic of the metric induced by the model's own learned potential — this is not a pure-conservation claim (the explicit \\(\gamma v\ell\\) damping term is included), only that the dynamics satisfy the damped geodesic equation with the architecture's own damping coefficient. Averaging over layers and 10 fixed validation batches (seed 42) gives \\(\bar{R}(\gamma{\text{train}})\\), evaluated at \\(\gamma{\text{eval}} = \gamma{\text{train}}\\) for each retained checkpoint — the diagonal overlay against \\(\mathrm{PPL}(\gamma_{\text{train}})\\).
How to read the scale — added 2026-09-26. The residual is normalised by the bare acceleration \\(\lVert a_\ell \rVert\\), which fixes three reference points:
This scale was not stated in the original version of this card, and it changes how the table below should be read. See Correction (2026-09-26) near the top.
A second, closed-form quantity — the recovered intrinsic damping \\(\gamma_{\text{geo}}\\) — is the least-squares row minimiser "the damping value that best explains this specific trajectory," independent of what \\(\gamma\\) the model was actually trained with:
$$ \gamma{\text{geo}} = -\frac{\big\langle a\ell + \Gamma(v\ell, v\ell),\ v\ell \big\rangle}{\lVert v\ell \rVert^{2}}. $$
Full derivation, practical mitigations (turning-point exclusion, reference-energy convention, integrator staggering), and validation controls (vanilla-baseline, shuffled-\\(\Gamma\\), and random-direction nulls) are in the companion note `Geodesic_Preservation_Experiment.md`.
Results: Minima Coincide at Gamma 0.10
<p align="center"><img src="results/geodesicoverlayanisogaussiand384.png" alt="PPL vs geodesic residual overlay for d=384 aniso-Gaussian + fock-reg gamma sweep" width="660"></p> The minima coincide. Both curves bottom out at \\(\gamma = 0.100\\) — a cleaner signal than the earlier \\(d=256\\) TinyStories aniso-Gaussian sweep, where the PPL-optimal (\\(0.150\\)) and \\(\bar{R}\\)-optimal (\\(0.050\\)) gammas disagreed. \\(\gamma_{\text{geo}}\\) also clusters tightly (\\(0.90\\)-\\(0.99\\) excluding the \\(\gamma=0.20\\) outlier) across the whole sweep, essentially independent of the nominal training \\(\gamma\\) — the same "intrinsic preferred geometry" signature documented at every other scale in this family.
The per-layer residual heatmap below shows where in the network the departures from geodesic behaviour concentrate at each damping level:
<p align="center"><img src="results/geodesicperlayeranisogaussian_d384.png" alt="Per-layer geodesic residual heatmap for d=384 aniso-Gaussian + fock-reg gamma sweep" width="660"></p>
The Gamma=0.20 Anomaly
Read in isolation, \\(\gamma=0.20\\)'s PPL of 2250 (8x worse than its neighbours) looks like a phase boundary. It is not: \\(\gamma=0.15\\) (PPL 283) and \\(\gamma=0.25\\) (PPL 292) bracket it and are both within 5% of the \\(\gamma=0.10\\) optimum, and a genuine stability wall would produce sustained degradation past the boundary rather than a single bad point sandwiched between two good ones.
The training log shows the signature of a real optimisation pathology, not measurement noise: train_ntp never dropped below ~7.5 (vs. 5.3-5.7 at neighbouring gammas) and the gradient norm was pegged at the per-group clip ceiling for nearly the entire run. val_ppl spiked to 9,257 at step 1,000 and 13,198 at step 1,500 before partially recovering to 2,250 by step 3,000 — a real divergence-and-partial-recovery excursion, most likely triggered by an early unlucky gradient/batch interaction the model never fully escaped within the 3K-step budget. Consistent with this, \\(\gamma_{\text{geo}} = 0.902\\) at \\(\gamma=0.20\\) is visibly lower than its neighbours (\\(0.982\\), \\(0.981\\)) — the divergent trajectory needed less retrofitted damping to explain its already highly non-geodesic path, exactly what a trajectory dominated by an early instability rather than smooth convergence would produce.
This checkpoint is retained in this repository specifically so this failure mode can be studied; it should not be read as a stability wall at \\(\gamma=0.20\\) for this architecture.
Coincidence or Crossover? Comparing Against the Isotropic Sibling at the Same Width
An earlier gamma sweep at this exact width and depth (\\(d=384\\), \\(L=16\\)) — using the family's original (non-anisotropic) \\(V_\theta\\) and no Fock-coupling regulariser, documented in the OpenWebText isotropic flagship's model card — found the opposite result: the PPL-optimal \\(\gamma\\) (0.25) and the geodesic-optimal \\(\gamma\\) (0.05) disagreed by a factor of 5, with the model preferring an overdamped, force-dominated regime that was measurably far from geodesic.
Two structural changes separate the two sweeps (anisotropic rank-4 precision instead of diagonal-only, and the Fock-coupling log-barrier regulariser instead of none), and this comparison cannot yet attribute the coincidence to one or the other in isolation. What it does establish is that the width-driven "phase transition" reported for the original \\(V\theta\\) family — overdamped-optimal below \\(d\approx768\\), near-geodesic-optimal above — is **not a fixed property of \\(d=384\\)**: a bounded, anisotropic \\(V\theta\\) with the coupling regulariser active pulls the low-\\(d\\) regime's optimum back toward the high-\\(d\\) anchor at this same width. This mirrors the \\(d=256\\) TinyStories finding (the low-\\(d\\) regime is present but dampened for bounded \\(V\theta\\), not eliminated outright) and is discussed at length, including the two-regime closed-form predictor's exact numeric comparison, in [§12 of `DeterminingoptimalgammaforFock-PARFLM.md`](https://github.com/dimitarpg13/semsimula-paper/blob/main/companionnotes/DeterminingoptimalgammaforFock-PARFLM.md#12-anisotropic-gaussian-v_theta-sweep-d384-l16-openwebtext-august-8-2026).
Caveats: Short-Sweep Reliability
This is a 3,000-step, single-seed measurement per gamma, and the family has one documented case where a short-sweep ranking reversed at full training length: the \\(d=256\\) TinyStories aniso-Gaussian sweep favoured \\(\gamma=0.150\\) at 3K steps, but the full 20K-step run favoured \\(\gamma=0.300\\) instead. Two considerations specific to this sweep are worth flagging before trusting \\(\gamma=0.10\\) at the full 100K-step horizon:
- \\(\gamma \in \{0.10, 0.15, 0.25\}\\) are all within 5% of each other in PPL — the same "flat bowl, ranking unreliable" signature that preceded the \\(d=256\\) reversal.
- The two-regime closed-form predictor (calibrated on the original, non-anisotropic \\(V_\theta\\) family) puts its low-\\(d\\) anchor at \\(\gamma \approx 0.246\\) and its high-\\(d\\) anchor at \\(\gamma \approx 0.050\\); the empirical \\(0.10\\) sits between them, closer to the high-\\(d\\) side but not an exact match to either.
It would not be surprising if the 100K-step full run ultimately favours something closer to \\(\gamma = 0.15\\)-\\(0.25\\) rather than this sweep's point estimate of \\(0.10\\). The full run's training log should be watched for the specific instability signature identified in the gamma=0.20 anomaly (gradient norm pegged near the per-group clip ceiling, train_ntp failing to drop into the 5.x range by step ~1,000) as an early-warning check that \\(\gamma=0.10\\) may be underdamped at this scale over a much longer horizon than the sweep tested.
How to Get Started
import math, torch, sys
sys.path.insert(0, "multixi")
sys.path.insert(0, "parf")
sys.path.insert(0, "energetic_minima")
sys.path.insert(0, "sarf_mass_variant")
from parf.model_fock_parf_multixi import FockMultiXiPARFLM, FockMultiXiPARFConfig
from parf.model_aniso_gaussian_vtheta import AnisotropicDepthConditionedGaussianVTheta, install_aniso_depth_routing
from huggingface_hub import hf_hub_download
REPO = "dimitarpg13/semsimula-fock-parflm-anisogaussian-vtheta-owt-d384-gammasweep"
GAMMA = "0.100" # the recommended candidate; also available: 0.050, 0.150, 0.200, 0.250, 0.300, 0.400, 0.500
logfreq_path = hf_hub_download(repo_id=REPO, filename="results/logfreq_surprisal_openwebtext.npy")
config = FockMultiXiPARFConfig(
vocab_size=50257, d=384, max_len=1024, L=16,
mass_mode="logfreq", logfreq_path=logfreq_path,
init_gamma=1.0, fixed_gamma=float(GAMMA),
xi_channels=5, xi_alpha_inits=[0.50, 0.75, 0.95, 0.99, 0.995],
xi_learnable=True, xi_alpha_init_mode="explicit",
fock_version="v2", n_registers=32,
reverse_channel=True, reverse_channel_stable=True, reverse_channel_pre_ln=True,
reverse_channel_soft_norm=True, reverse_channel_warmup_steps=4000, reverse_channel_per_layer=True,
register_repulsion=True, register_repulsion_coeff=0.05,
prefix_causal_registers=True,
v_phi_kind="structural_competitive", v_phi_n_heads=4, v_phi_d_type=32, v_phi_d_angle=16, top_k=16,
use_output_bias=True, tie_embeddings=False,
per_register_tau=True, per_register_keys=True, ortho_register_init=True,
)
model = FockMultiXiPARFLM(config)
model.V_theta = AnisotropicDepthConditionedGaussianVTheta(
d=384, K=8, n_ctx=5, n_layers=16, rank=4,
w_scale=1.0,
init_log_precision=-math.log(384),
precision_max=2.0 / 384,
code_init_std=0.02,
)
install_aniso_depth_routing(model)
ckpt_path = hf_hub_download(repo_id=REPO, filename=f"checkpoints/gamma_{GAMMA}/ckpt_best.pt")
state = torch.load(ckpt_path, map_location="cpu", weights_only=False)
model.load_state_dict(state["model_state_dict"])
model.eval()
print(f"Parameters: {sum(p.numel() for p in model.parameters()):,}") # 77,012,107
print(f"gamma={state['gamma']} step={state['step']:,} val_ppl={state['val_ppl']:.2f}")Available Artifacts
Training Details
Training Data
OpenWebText, tokenized with GPT-2 BPE (vocab 50257). Each candidate trains on up to 1B tokens (early stopped at 3,000 steps, effective batch 16, block 512 -> ~24.6M tokens actually consumed) and is evaluated on a held-out 2M-token validation slice with no train/val overlap.
Training Procedure (per gamma candidate)
Causal-Leak Verification
All 8 checkpoints were trained natively with prefix_causal_registers=True from step 0. The bit-exact future-perturbation causal probe passed with max_delta=0.0 at step 2,000 for every one of the 8 gamma candidates, including the \\(\gamma=0.20\\) divergence case — the instability there is an optimisation pathology, not a causal-leak symptom.
Training Script
notebooks/conservative_arch/scaleup/colab_fock_gamma_sweep_geodesic_aniso_gaussian_fockreg_d384.ipynb (companion repo) — self-contained Colab notebook that runs the 8-candidate sweep, the geodesic residual analysis, and produces the overlay/heatmap figures in one pass, requiring no additional training beyond the sweep itself.
Evaluation Results
OpenWebText Validation Perplexity (3,000-step sweep candidates)
These PPL values are from 3,000-step short-sweep candidates and are not comparable to the fully trained OpenWebText checkpoints elsewhere in this family (e.g. 27.23 PPL after 250K steps for the isotropic flagship). They exist solely to rank candidate damping coefficients.
See Results: Minima Coincide at Gamma 0.10 for the combined PPL / geodesic-residual table.
SPLM Family Overview
This model is part of the Semantic Simulation SPLM family:
Collection: Semantic Simulation SPLM Model Family
Bias, Risks, and Limitations
- Not a final trained model. Every checkpoint in this repository has been trained for only 3,000 steps — a gamma-selection diagnostic, not a language model intended for generation or downstream use. Do not compare its PPL to fully trained checkpoints elsewhere in this family.
- Single-seed sweep. Each gamma candidate is one run; no seed-variance estimate is available. The flat-bowl caveat in Caveats: Short-Sweep Reliability applies.
- Short-horizon ranking can reverse. The family has one documented case (\\(d=256\\) TinyStories) where the short-sweep-optimal gamma differed from the full-run-optimal gamma. The \\(\gamma=0.10\\) recommendation here carries the same risk; see the caveats section.
- The gamma=0.20 checkpoint is a divergence artefact, not a representative trained model at that damping value; see The Gamma=0.20 Anomaly.
- OpenWebText only, English only. No instruction tuning, no RLHF/DPO, no safety filtering.
- Geodesic residual is a diagnostic, not a training objective. The model was trained to minimise cross-entropy; the near-geodesic behaviour at \\(\gamma=0.10\\) is an emergent structural property, not something the loss function directly optimises for.
- Large \\(V_\theta\\) hypernetwork. At 35.5M parameters, \\(V_\theta\\) alone is nearly half of the 77M-parameter total — substantially larger (in absolute terms and in share of the model) than any prior sibling in this family, driven by the width increase (\\(d=384\\)) and the 5-context, rank-4 low-rank correction.
- No causal-leak issue. All 8 checkpoints were trained natively with
prefix_causal_registers=True; see Causal-Leak Verification.
Citation
@misc{Gueorguiev2026SemSim,
author = {Gueorguiev, Dimitar P.},
title = {Semantic Simulation: A Prescriptive Lagrangian Framework
for Efficient Semantic Inference --- A Conservative-by-
Construction Language Model and the Shared-Potential
Separator, with a Correspondence to Joint Embedding
Predictive Architectures},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.19712427},
url = {https://doi.org/10.5281/zenodo.19712427},
note = {Companion code repository:
\url{https://github.com/dimitarpg13/semsimula-paper}}
}Environmental Impact
- Hardware: 1x NVIDIA H100/A100 80GB (Google Colab)
- Training: 8 x 3,000 steps = 24,000 total training steps across the sweep, plus inference-only geodesic residual analysis (10 validation batches x 8 checkpoints, no additional training)
- Carbon footprint: small; a single-GPU research sweep, estimated on the order of a few kg CO2
Correction (2026-09-27) — integrator name. This model's layer step was described here as a damped Euler step. It is not. The update ish_new = h + (h - h_prev)/(1 + dt*gamma) + dt^2*f/(m*(1 + dt*gamma)), which carries no velocity state at all: the velocity is the position differenceh - h_prev. Undamped this ish_{n+1} = 2*h_n - h_{n-1} + dt^2*f/m, i.e. Störmer–Verlet in position form. It is also not velocity-Verlet, which carries an explicit velocity through half-kick/drift/half-kick. The distinction matters because the \\(\omega \cdot dt < 2\\) stability wall is the Störmer/leapfrog bound and applies to this lineage; the semi-implicit Euler models in the same collection were never subject to it. Only the name was wrong — no measurement on this card changes.
