CoolFace
Modelpublic

knightnemo/wuji-hand-gesture-vam-ti2v5b-30l-openwam-fastwam-f33-r4-h50-ref65-refdrop10-mask-v3-propdrop10

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes
Model Card

Wuji Hand Gesture Vam Ti2V5B 30L Openwam Fastwam F33 R4 H50 Ref65 Refdrop10 Mask V3 Propdrop10

This repository contains one Wuji hand gesture VAM checkpoint from the May 26, 2026 OpenWAM/FastWAM sweep.

Identity

text
repo_id:              knightnemo/wuji-hand-gesture-vam-ti2v5b-30l-openwam-fastwam-f33-r4-h50-ref65-refdrop10-mask-v3-propdrop10
wandb project:        wuji_hand_gesture
wandb run id:         vtyfti4s
wandb run name:       wuji_hand_gesture_vam_ti2v5b_30L_mask_v3_openwam_refdrop_0.1_propdrop_0.1_0526_1545
local training dir:   /cpfs/huangsq/VAM_Learn_from_Human_Video/src/vam/models/train/wuji_hand_gesture_vam_ti2v5b_30L_openwam_fastwam_f33_r4_h50_ref65_refdrop10_mask_v3_propdrop10
checkpoint:           step-10000.safetensors
checkpoint size:      12041754617 bytes
base model:           Wan-AI/Wan2.2-TI2V-5B
action expert style:  openwam
mask variant:         v3
action horizon:       50
proprio dropout:      0.1
reference dropout:    0.1

The checkpoint is a joint model: step-10000.safetensors contains the fine-tuned video DiT weights plus the action_dit.* action-stream weights. The Wan2.2-TI2V-5B base model is not included.

Files

text
step-10000.safetensors   final 10k-step checkpoint
model_config.json        compact machine-readable configuration and metrics
training_config.yaml     full W&B training config snapshot
wandb-summary.json       final scalar metrics exported by W&B
training_log_node0.txt   node-0 training log
README.md                this model card

Final Step Metrics

These are the scalar values in wandb-summary.json at step=10000.

MetricValue
val/loss0.303067
val/loss_action0.245217
val/loss_video0.0578496
val/action_mse0.0686766
val/action_mae0.166933
val/video_mse162.44
val/video_psnr29.5267
val/video_ssim0.956715
val/video_lpips0.0371108
train/loss0.00450895
train/loss_action0.000255521
train/loss_video0.00425343
train/grad_norm0.490234
_runtime15233.5

Best saved checkpoint by aggregate validation loss during the run:

text
step 5000: loss=0.208117, loss_video=0.068828, loss_action=0.139289

Only step-10000.safetensors is uploaded here, so use the best-saved entry only as training provenance unless that step is also uploaded separately.

Training Configuration

KeyValue
dataset_typewuji_gesture
wuji_robot_dataset_root/cpfs/huangsq/wuji-hand-gestures-cropped
variantclean_50
height256
width256
num_frames33
action_video_freq_ratio4
action_horizon50
action_dim20
proprio_dim20
action_formatabsolute
action_spacejoint
proprio_spaceaction
action_pad_modelast
action_expert_styleopenwam
action_mot_backbone_pretrained_path/cpfs/huangsq/VAM_Learn_from_Human_Video/src/vam/models/pretrained/ActionMoT_openwam_linear_interp_Wan22_alphascale_1024hdim.pt
mask_variantv3
mask_tail_padding_lossTrue
full_reference_videoTrue
max_ref_frames65
reference_dropout0.1
bridge_exclude_full_refTrue
extra_inputsvace_reference_image,action_trajectory
target_camerahead_camera
reference_camerahead_camera
resize_modepad
backboneti2v
model_paths["models/Wan-AI/Wan2.2-TI2V-5B/diffusion_pytorch_model-00001-of-00003.safetensors","models/Wan-AI/Wan2.2-TI2V-5B/diffusion_pytorch_model-00002-of-00003.safetensors","models/Wan-AI/Wan2.2-TI2V-5B/diffusion_pytorch_model-00003-of-00003.safetensors","models/Wan-AI/Wan2.2-TI2V-5B/models_t5_umt5-xxl-enc-bf16.pth","models/Wan-AI/Wan2.2-TI2V-5B/Wan2.2_VAE.pth"]
tokenizer_pathmodels/Wan-AI/Wan2.2-TI2V-5B/google/umt5-xxl
trainable_modelsdit
learning_rate5e-05
action_lrNone
weight_decay0.01
warmup_steps500
max_steps10000
num_epochs1
batch_size1
gradient_accumulation_steps1
dataset_repeat1
dataset_num_workers8
use_gradient_checkpointingTrue
save_steps1000
val_steps200
video_log_steps1000
max_val_samples20
lambda_video1
lambda_action1
video_dim3072
action_dit_dim1024
action_dit_ffn_dim4096
action_dit_num_heads24
action_dit_num_layers30
proprio_dropout0.1
window_stride1
val_ratio0.1
output_path/cpfs/huangsq/VAM_Learn_from_Human_Video/src/vam/models/train/wuji_hand_gesture_vam_ti2v5b_30L_openwam_fastwam_f33_r4_h50_ref65_refdrop10_mask_v3_propdrop10
wandb_projectwuji_hand_gesture
wandb_run_namewuji_hand_gesture_vam_ti2v5b_30L_mask_v3_openwam_refdrop_0.1_propdrop_0.1_0526_1545

Input/Output Contract

Expected inputs:

text
prompt:                 "the robot performs hand gesture {label}"
target camera:          head_camera
reference camera:       head_camera
target video frames:    33
full reference frames:  65
image resolution:       256 x 256
action horizon:         50
action/proprio dim:     20 / 20

Expected outputs:

text
robot-view target video rollout
20-D absolute robot action targets

Masking Note

This run uses mask_variant=v3 with the sequence layout:

text
[ref_video | first_frame | gen_video | action]

The run also records bridge_exclude_full_ref=True for provenance. For this OpenWAM/ActionMoT path, the active v2/v3 distinction is the mask_variant listed above.