abdullah-hmed/vjepa_21_t2i_adapter_sd15
A T2I Adapter trained on VJEPA 2.1 embeddings, to reconstruct the embeddings back into pixel space.
Some Examples:
P.S. Most of these examples were generated using the LCM Scheduler + LCM LoRA for 2-4 step per frame generations. Using DDIM would yield much better results, and generate much more faithful colors
<video src="https://huggingface.co/abdullah-hmed/vjepa21t2iadaptersd15/resolve/main/examples/inputcomparison20260513_095924.mp4" controls autoplay loop width="600"></video>
<video src="https://huggingface.co/abdullah-hmed/vjepa21t2iadaptersd15/resolve/main/examples/birbcomparison20260513_022009.mp4" controls autoplay loop width="600"></video>
Even facial details are conserved if the face occupies a significant part of the frame
<img src="https://huggingface.co/abdullah-hmed/vjepa21t2iadaptersd15/resolve/main/examples/balecomparison20260514_111333.png" width="600"/>
<img src="https://huggingface.co/abdullah-hmed/vjepa21t2iadaptersd15/resolve/main/examples/IMG9066comparison20260515033902.png" width="600"/>
