CoolFace
Modelpublic

SepehrNoey/MCM-Simplified

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes12downloads
Model Card

MCM-Simplified

<Gallery />

Model description


license: apache-2.0 tags:

  • text-to-video
  • motion-consistency
  • distillation ---

Motion Consistency Model - Simplified Implementation

This model is a distilled version of the Motion Consistency Model, trained on a subset of WebVid 2M with additional filtered image-caption pairs from the LAION aesthetic dataset.

Sample Generated Videos

CaptionTeacher (ModelScope) - 50 DDIM StepsStudent (First Setup) - 4 StepsStudent (Second Setup) - 4 Steps
Worker slicing a piece of meat.ImageImageImage
Pancakes with chocolate syrup, nuts, and bananas.ImageImageImage

Training Details

  • Dataset: 3022 video-caption pairs from WebVid 2M
  • Image Pairs:
  • Setup 1: 20K filtered LAION aesthetic images (min. resolution 450×450)
  • Setup 2: 7.5K filtered LAION aesthetic images (min. resolution 1024×1024)

Training Configurations

Setup 1
  • LR: 5e-6, Grad Accum: 4, Max Grad Norm: 10
  • Discriminator LR: 5e-5, Weight: 1, Lambda R1: 1e-5
  • EMA Decay: 0.95, Epochs: 7, Steps: ~5100
Setup 2 (Modified)
  • LR: 2e-6, Grad Accum: 16, Max Grad Norm: 5
  • Discriminator LR: 1e-6, Weight: 0.5, Lambda R1: 1e-4
  • EMA Decay: 0.98, LR Warmup: 300 steps, Epochs: 10

Evaluation

Frechet Video Distance (FVD)

Model1 Step2 Steps4 Steps8 Steps
Teacher (50 DDIM Steps)2954.77---
Student - Setup 12598.152684.243082.843914.78
Student - Setup 22589.013053.353284.693930.07

CLIP Similarity (×100)

Model1 Step2 Steps4 Steps8 Steps
Teacher (50 DDIM Steps)27.88---
Student - Setup 122.5525.6226.8627.01
Student - Setup 220.1323.4125.3124.62

Conclusion

Setup 2 was modified to stabilize training and prevent the discriminator from overpowering the generator. The changes improved FVD scores for 1-step inference, while multi-step performance varied. CLIP similarity improved across multiple inference steps, indicating better text-to-video alignment.

References

Original Implementation: Motion Consistency Model

Download model

Weights for this model are available in Safetensors format.

Download them in the Files & versions tab.