yxma/vtwm-ablation-comparison
0
vm_diffusion rollouts
Diffusion-forcing world model rollouts. Each sample_XXX/ is a window from the validation set: ~2.7 s of video. The model predicts the future half of each window given the past half as context.
Vision-only models still emit zero tactile latents (so the tactile videos in those runs will look like uniform gray noise — that's expected).
