Lawrence-cj/h3-ltx25-latent-adapter
H3 → LTX-2.5 latent adapter showcase
This static Hugging Face Space presents the frozen exact-32 evaluation of the MiniMax H3 → LTX-2.5 Conv Video VAE latent adapter. It contains pre-rendered videos and does not require a GPU or run the VAEs in the browser.
Scope
- Validated target: LTX-2.5 Conv Video VAE latent space.
- Weight identity proof: the local VAE checkpoint used by every paired target, training run, and decode has SHA-256
425f0dfa227dee5d0ff3d9720563370810409a439c302ca74f0f944057ce55c5. That exactly matches the officialLightricks/LTX-2.5-Diffusers/vae/diffusion_pytorch_model.safetensorsobject (1,452,233,194 bytes). - The main showcase uses the faster convolutional decoder. A separate
benchmark.htmlpage visually compares it with a real three-step LTX-2.5 distilled-transformer latent Refiner on 16 frozen validation clips. multires.htmlkeeps the direct Adapter PSNR/SSIM table for all eight fixed clips at each of six resolutions, and now shows two fixed visual examples per resolution as H3 source | Adapter direct decode | Adapter + real three-step Refiner. Refiner visuals are intentionally not assigned PSNR/SSIM.two_stage.htmlevaluates the official LTX-2.5 two-stage geometry on eight motion-stratified clips: 544×960 Adapter latent → official x2 spatial latent upsampler → 1088×1920 three-step Refiner → Conv Video VAE. Each visual keeps the source, 544p direct decode, upsampler-only decode, and refined decode; Refiner PSNR/SSIM remains intentionally absent.- Displayed deployment pipeline: existing H3 VAE latent → deterministic geometry packing → adapter → frozen LTX-2.5 Conv Video VAE decoder. The H3 source and LTX teacher reconstruction are evaluation references only; an H3 encoder is not required when the latent is already available.
- No LTX diffusion transformer or latent upsampler is used for the main Adapter result. The separate Refiner benchmark starts from the exact same predicted LTX latent, applies Gaussian noise and three transformer Euler updates, then decodes both paths with the same Conv VAE.
Publish
Create a new Hugging Face Space with the Static SDK, then push the contents of this directory to the root of the Space repository. The YAML block above is the complete Space configuration; index.html is the app entry point.
The bundled MP4 files are web display encodes, so the current showcase does not require Git LFS.
Source artifacts
The two triptych videos come from the promoted checkpoint's exact-32, 192-frame, 768×1344 geometry-audited evaluation. Each triptych is ordered:
- source H3-generated video,
- frozen LTX-2.5 Conv Video VAE teacher reconstruction,
- adapter-predicted LTX latent decoded by the same frozen LTX-2.5 Conv Video VAE.
Promoted checkpoint: output/h3_ltx_latent_adapter/runs/formal_leader_plus_native16384_minus_native8192/checkpoint.pt. It is a 194,759,504-parameter, 22-block factorized Conv3D adapter; no Transformer is present in the adapter.
The mean promoted result is 28.77923 dB PSNR and 0.87486 SSIM against the LTX VAE teacher over the frozen 32-video validation set. Mean temporal-gradient MAE is 0.02020 and the remaining gap to 29 dB is 0.22077 dB.
For the direct end-to-end path requested here, the adapter decode is 28.71090 dB PSNR / 0.85911 SSIM against the original H3-generated source. The frozen LTX teacher reconstruction itself is 34.99138 dB / 0.94715 against that same source. The page keeps all three comparisons visible so the cross-VAE adapter error is not confused with the LTX VAE reconstruction floor.
The 78,400-pair standard direction adds 0.08061 dB over the previous leader; calibrating that direction to 3.00x reaches 28.71754 dB. The direct 3.00x versus 2.00x gain is 0.04270 dB with paired 95% CI [0.03781, 0.04749]; PSNR and SSIM improve on all 32 validation IDs. Replacing the older first-layer sensitivity with the pooled full-decoder-Jacobian direction adds another 0.01035 dB with paired 95% CI [0.00728, 0.01358], reaching its registered intermediate leader. The native-8K matched-exposure direction then adds 0.01249 dB on exact-32 with a positive paired interval and PSNR gains on all 32 IDs. The bounded spatial decoder-crop branch is closed because no halo through 12 passed the fixed worst-path fidelity gate.
The preregistered direction-scale grid is closed. Scale 4 was the only development candidate with a positive paired PSNR interval and subsequently passed both disjoint audit and original exact-32. It adds +0.02127 dB over the previous leader with 95% CI [+0.01610, +0.02640]. The subsequent stable 128×128 decoder-Jacobian channel-Gram replacement passes fresh development, disjoint audit, and formal exact-32. It adds +0.00486 dB over scale 4 with 95% CI [+0.00421, +0.00556] and improves PSNR and SSIM on all 32 formal IDs, without changing the adapter's inference architecture or parameter count.
The subsequent fixed decoder-frequency replacement also passes fresh development, disjoint audit, and formal exact-32 gates. It adds +0.00133 dB over the Gram leader with paired 95% CI [+0.00112, +0.00153] and 31/32 PSNR wins. Mean SSIM adds +0.00001165 with a positive interval, and temporal error falls on all 32 IDs. The ConvNeXt change is unresolved. Frequency weighting is used only during training, so inference architecture and parameter count stay unchanged.
The native-16K matched-exposure direction is now promoted. It passes untouched development (+0.01392 dB), disjoint audit (+0.01579 dB), and original formal exact-32 (+0.01139 dB) with positive paired PSNR intervals and nonnegative mean SSIM at every gate. The formal 95% PSNR interval is [+0.00851, +0.01412] dB; formal SSIM adds +0.000461. The direction is native16K-step3200 - native8K-step1600, composed once onto the frequency leader without changing inference size.
Current work is not included in the promoted headline metrics. The preregistered strong-parent consolidation starts from this exact leader with a fresh optimizer. Its fixed recipe is 3,200 steps on 16,384 native pairs with latent MSE + 160× frozen-LTX-teacher pixel MSE. The first 800-step block and its endpoint audit pass. Internal validation changes by +0.06177 dB PSNR and +0.00146 SSIM, but those numbers are diagnostic only. The remaining blocks resume through a lock-protected one-shot coordinator when the scheduler is available. It permits one controller query only at a four-hour boundary and creates or adopts the complete audited 800→1,600→2,400→3,200 chain together with native-16K motion scoring. Only the exact step-3,200 endpoint may enter fresh development, disjoint audit, and formal exact-32 gates.
If that candidate promotes at or below 30 dB, its registered 3,200→16,384 unique-data continuation is now execution-complete but still inactive. Seventeen immutable endpoints cover steps 4,000, 4,800, ..., 16,000, and 16,384. Every block freezes checkpoint/config/metrics, proves continuous steps and a no-repeat seeded permutation, and resumes only through the prior endpoint audit. The launcher adopts only a complete audited prefix and makes one queue query; readiness is currently absent, so it has submitted no job and produced no model result.
Its future evaluation inputs are now prepared without exposing them early. Only after readiness, seeds 20261225/20261226 select a new validation32 and a disjoint test32 after excluding 22 prior manifests and all historically decoded IDs. H3 and LTX encodes then run in parallel per split, and each cache is accepted only after an exact 192×768×1344 metadata/shape rescan. The current manifests and caches do not exist; this is workflow readiness, not evidence of quality improvement.
The full-exposure selection path is also implementation-complete. It advances development, disjoint audit, and formal exact-32 progressively with frozen bootstrap seeds and joint PSNR/SSIM plus Q4 motion-detail gates. A formal pass must produce paired parent/candidate renders for the exact eight highest-motion IDs; no new ghosting, action blur, or color instability is allowed. Only then can the fail-closed resolver select a leader and evaluate the strict >30 dB GAN phase boundary. No full-exposure decode or result exists yet.
The activation cutoff was amended from the earlier 29 dB milestone to the current distortion target before consolidation quality was observed. A promoted result not strictly above 30 dB must therefore complete the same unique-exposure treatment before motion-loss parent selection.
The eligible 11:09:31 PT quality-coordinator call found the Slurm controller unavailable and submitted zero jobs. No model or metric changed; the next eligible recovery boundary is 15:09:31 PT.
GAN enhancement is a later phase, not part of the current optimization. It is forbidden until formal exact-32 mean PSNR is strictly greater than 30.0 dB at 192×768×1344. The current leader is 1.22077 dB below that boundary. Development/audit scores and individual sample maxima do not count. If the gate is eventually crossed, adversarial training must be a new matched fork from the immutable PSNR leader; distortion and perceptual leaders remain separate.
Fast-motion quality plan
Fast-action blur is now treated as a separate failure mode instead of being inferred from aggregate PSNR. The evaluator freezes motion masks and backward Farneback flow from the LTX teacher, ranks the exact 32 IDs into four motion quartiles, and reports motion-masked PSNR/SSIM, moving-edge error and sharpness, temporal velocity and acceleration error, and flow-warp residual.
A candidate-independent source-motion diagnostic is available for the current formal leader. On the same exact 32 IDs, prediction-vs-teacher PSNR changes from 31.3909 dB in the lowest-motion quartile to 26.1076 dB in the highest-motion quartile; temporal-gradient MAE changes from 0.00863 to 0.03255. The LTX teacher reconstruction is itself 3.3627 dB harder in the highest quartile, but the additional cross-VAE PSNR penalty still increases by 1.9258 dB with unpaired bootstrap 95% CI [+0.8112, +3.1089]. Different IDs occupy the two quartiles, so content difficulty remains a confounder and this diagnostic cannot select or promote a checkpoint. It establishes temporal capacity and motion-aware supervision as the highest-priority mechanisms to test.
The two already-exported triptychs provide a narrower render-level check. In the available Q4 example, the prediction retains 67.7% of teacher edge energy and 53.6% of teacher Laplacian energy inside the top-20% teacher-motion mask; the available Q2 example retains 74.5% and 68.1%. A bounded ±2-frame and ±2-pixel search selects zero lag and zero translation for both. This supports high-frequency attenuation rather than a simple global timing offset, but two downsampled H.264 renders are exploratory mechanism evidence, not exact-32 coverage or checkpoint-selection evidence.
That observation now has a preregistered exact-32 safeguard. Fresh motion decodes add teacher-moving-edge-weighted Laplacian-magnitude L1 and the absolute log ratio of prediction/teacher Laplacian energy on the existing 16-frame 192×336 auxiliary grid, while the frozen VAE still decodes the complete 192×768×1344 video. In Q4, both reductions must be nonnegative against control and parent. The gate requires an exact field set, so legacy metrics cannot be adopted or silently backfilled. This is evaluation logic only, not a new model result.
The matched training treatment now mirrors that safeguard. Control remains latent MSE + 160 × full-grid teacher pixel MSE; treatment adds teacher-motion-weighted edge, Laplacian-magnitude, velocity, and acceleration Charbonnier losses on the 16-frame 192×336 auxiliary grid. Eight calibration clips (two per frozen motion quartile) set each weight from its worst relative parameter-gradient norm. Every term receives a nominal 2.5% share and their combined per-sample upper bound cannot exceed 10% of weighted pixel MSE. This adds no inference parameter and is not yet a quality result.
Historical exact-32 evidence closes two tempting detours. The current linear-plus-nearest-pack representation preserves all 57 H3 tokens once, so the linear anchor does not discard the original temporal samples. A matched coordinate-10x arm adds only +0.00017 dB overall (confidence interval crosses zero) and +0.00019 dB in Q4, while raw MSE already beats log-MSE by +0.06842 dB overall and +0.08052 dB in Q4. GPU priority therefore remains on the four-term motion loss and then unique high-motion data, not another timing metadata stem or log-MSE rerun.
On these 32 IDs, source motion is not positively correlated with the recorded scene-cut ratio (-0.245). Its association with the adapter-specific PSNR penalty remains +0.604 after controls for teacher reconstruction difficulty, scene cuts, global luma, and edge density (bootstrap 95% interval +0.343 to +0.794). This is an observational diagnostic only, not a causal or selection claim.
Historical matched attribution does not support dense channel-time mixing as the first response. On the same exact32 geometry, the older dense-temporal arm is -0.00375 dB versus its depthwise control (95% CI -0.00452 to -0.00299), and Q4 is -0.00354 dB (95% CI -0.00504 to -0.00222). Because that run used a latent-only objective and a higher learning rate, the new decoder-aware dense-temporal canary remains prepared, but only as a fallback after the full-grid motion-loss A/B.
The native-16K source-motion preprocessing stage is ready for reliable recovery. Its 64 CPU shards can be resubmitted after partial completion: existing output is adopted only when its exact manifest partition, source identity, geometry, video hashes, finite diagnostics, and unique IDs all match. Merge requires exactly shard indices 0–63. This is pipeline readiness, not a model result; no scoring or training job has been launched by the change.
A single pinned launcher now owns the score-to-merge dependency chain. It can adopt one matching active score array and merge, rejects ambiguous jobs or a wrong merge dependency, and validates canonical merged outputs before a no-op. This removes manual dependency wiring; it does not change the current model or headline metrics.
The GPU motion-loss experiment now also has one fail-closed launcher. It recomputes the final parent from the live three-gate consolidation resolution, validates exact native-16K scores, and submits calibration → materialization → matched control/treatment step 400. Existing calibration or config artifacts are adopted only after complete identity, gradient, memory, and readiness revalidation; partial output is rejected. This remains execution status, not a new model result.
Its post-step-400 evaluation is now one pinned dependency chain as well. Both arms are audited at the exact endpoint, then the parent, control, and treatment are decoded on the same validation exact32 geometry. Treatment must pass two separate standard-plus-high-motion gates. Existing comparison and gate files are accepted only after exact recomputation from live metrics, including the registered 10,000 bootstrap samples; a partial artifact set fails closed. Only both passing gates can create 400→800 resume readiness, and step 400 cannot promote. All 380 adapter tests pass. The one 06:00 PT controller check found Slurm unavailable, so the experiment remains unlaunched and headline quality is unchanged.
The fallback architecture A/B changes only the last three temporal blocks. Its control retains depthwise temporal convolutions and trains 9,024 parameters. Its treatment embeds those kernels on the diagonal of dense channel-time convolutions, is exactly function preserving at initialization, and trains 5,091,792 parameters inside a 199,842,272-parameter model. Both arms otherwise share the parent, native-16K order, seed, 1e-6 learning rate, and latent MSE + 160 × pixel MSE objective.
The first post-parent mechanism A/B keeps the frozen LTX VAE decode at the complete 192×768×1344 grid. It downsamples only the auxiliary loss graph to 192×336 and adds teacher-motion-weighted moving-edge, Laplacian-magnitude, velocity, and acceleration Charbonnier terms. Their combined parameter-gradient norm is capped at 10% of the weighted pixel-MSE gradient. A cropped decoder treatment was rejected because every tested halo from 0 through 12 failed the fixed full-decoder fidelity gate.
The execution path is now fail-closed end to end. Before either arm can start, one readiness artifact rechecks the resolved parent SHA, exact 16,384 source- motion scores with 4,096 clips per quartile, balanced worst-sample calibration, and the real memory gate. Both arms stop at step 400. They can resume together to step 800 only after treatment beats both control and parent on the joint development PSNR/SSIM and highest-motion gates. This is implementation status, not a new model result.
Neither experiment has a promoted result yet. Step 400 is only a continuation gate: the treatment must beat its matched control and the resolved parent on development. Step 800 must then pass fresh development, disjoint audit, and formal exact-32 paired PSNR/SSIM gates, all registered highest-motion no-regression checks, and visual review of the eight highest-motion videos for new ghosting, blur, or color instability. The Space headline changes only after that complete evidence chain passes.
The next data A/B, if needed, will not repeat high-motion clips. Both arms use 800 unique native-16K IDs. Control uses 25% from every motion quartile; treatment uses the registered 2:3:4:5 Q1-Q4 allocation, approximately 14.3%/21.4%/28.6%/35.7%, with exact 57/86/114/143 counts in each 400-step block. The model, objective, optimizer, and update count stay matched. The uncalibrated 50%-Q4 proposal is superseded. This branch is implemented but has no selection authority and remains inactive until the first motion-loss experiment resolves; fresh development and audit evidence is mandatory.
A separate source-detail capacity fallback is now prepared below those two priorities. A zero-initialized branch reads aligned H3 latents through a 1x1 stem and four width-256 factorized Conv3D blocks. Its raw-source control and fixed temporal-spatial-high-pass treatment share identical new tensors and train 4.076M parameters each. The real formal checkpoint expands from 194.759M to 198.836M parameters with exact-zero output change. Both arms use the same native-16K full-grid decoder objective and stop at step 400 for the existing dual standard/Q4 gate. All 390 tests pass, but no job or quality result exists; the branch runs only if motion loss and calibrated unique-data exposure leave a residual fast-motion failure.
Scheduler recovery is intentionally low frequency. The one-shot prerequisite coordinator uses a durable four-hour cadence, a process lock, and exactly one controller query per eligible invocation. A fresh recovery simulation creates eight consolidation jobs and two native-16K motion score/merge jobs with exact afterok dependencies; controller-down, within-window, ambiguous-job, and dependency-drift cases all fail without submission. The coordinator closed at 397 tests; the complete suite now passes 403 tests after the render diagnostic and CLI-safety amendment. This is workflow evidence only: it has not changed the checkpoint, metrics, or public Space revision.
One operational correction is recorded explicitly. An attempted --preflight at 07:09 PT exposed that the original shell entrypoint ignored positional arguments and performed one real controller query. The controller was still unavailable and no job was submitted. The entrypoint now supports that offline flag and rejects unknown arguments before scheduler access. Its durable state sets the next eligible check to 11:09:31 PT; routine execution must not bypass that cooldown.
