CoolFace
Apppublic

WitneyWW/vmx-injection-gt-vs-pred-strips

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes
App README

vmx injection - ground truth vs prediction strips

GT-over-prediction frame strips for the frozen view-motion-encoder (vmx) injection into the Video+Tactile Mixture-of-Transformers diffusion-forcing world model.

Each strip shows, per stream (camera view, tactile left, tactile right): row 1 = ground truth, row 2 = prediction. A red divider marks where prediction begins; frames left of it are given context.

Four cells per run: {short window, long 16s rollout} x {test, train}, 10 rows each. Every start index leaves at least 96 frames of real ground truth inside the episode, so the long cells are never compared against clamped tail frames.

Test episodes (motherboard_0510_episode_005/006) are held out of training; train rows are from episodes the model saw.