abdullah-hmed/streetwm-vjepa21
0
An attempt at training a video continuation model by training a predictor on optical flow, and V-JEPA 2.1 encoded frame embeddings of an approx. 1hr 30m sidewalk traversal video
Looks like this: <video src="https://huggingface.co/abdullah-hmed/streetwm-vjepa21/resolve/main/worldmodel20260428-232050.mp4" controls autoplay loop/>
