CoolFace
Modelpublic

knightnemo/wuji-hand-gesture-vam-ti2v5b-30l-openwam-fastwam-f33-r4-h32-ref65-refdrop10-mask-v2-propdrop30

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes
Model Card

Wuji Hand Gesture Vam Ti2V5B 30L Openwam Fastwam F33 R4 H32 Ref65 Refdrop10 Mask V2 Propdrop30

This repository contains one Wuji hand gesture VAM checkpoint from the May 26, 2026 OpenWAM/FastWAM sweep.

Identity

text
repo_id:              knightnemo/wuji-hand-gesture-vam-ti2v5b-30l-openwam-fastwam-f33-r4-h32-ref65-refdrop10-mask-v2-propdrop30
wandb project:        wuji_hand_gesture
wandb run id:         bqqhtqyy
wandb run name:       wuji_hand_gesture_vam_ti2v5b_30L_mask_v2_openwam_refdrop_0.1_propdrop_0.3_0526_1545
local training dir:   /cpfs/huangsq/VAM_Learn_from_Human_Video/src/vam/models/train/wuji_hand_gesture_vam_ti2v5b_30L_openwam_fastwam_f33_r4_h32_ref65_refdrop10_mask_v2_propdrop30
checkpoint:           step-10000.safetensors
checkpoint size:      12041754617 bytes
base model:           Wan-AI/Wan2.2-TI2V-5B
action expert style:  openwam
mask variant:         v2
action horizon:       32
proprio dropout:      0.3
reference dropout:    0.1

The checkpoint is a joint model: step-10000.safetensors contains the fine-tuned video DiT weights plus the action_dit.* action-stream weights. The Wan2.2-TI2V-5B base model is not included.

Files

text
step-10000.safetensors   final 10k-step checkpoint
model_config.json        compact machine-readable configuration and metrics
training_config.yaml     full W&B training config snapshot
wandb-summary.json       final scalar metrics exported by W&B
training_log_node0.txt   node-0 training log
README.md                this model card

Final Step Metrics

These are the scalar values in wandb-summary.json at step=10000.

MetricValue
val/loss0.291118
val/loss_action0.229092
val/loss_video0.0620255
val/action_mse0.046491
val/action_mae0.143964
val/video_mse87.1846
val/video_psnr30.4364
val/video_ssim0.962145
val/video_lpips0.0220791
train/loss0.0308171
train/loss_action0.0130898
train/loss_video0.0177273
train/grad_norm0.886719
_runtime15291.6

Best saved checkpoint by aggregate validation loss during the run:

text
step 7000: loss=0.230949, loss_video=0.060932, loss_action=0.170017

Only step-10000.safetensors is uploaded here, so use the best-saved entry only as training provenance unless that step is also uploaded separately.

Training Configuration

KeyValue
dataset_typewuji_gesture
wuji_robot_dataset_root/cpfs/huangsq/wuji-hand-gestures-cropped
variantclean_50
height256
width256
num_frames33
action_video_freq_ratio4
action_horizon32
action_dim20
proprio_dim20
action_formatabsolute
action_spacejoint
proprio_spaceaction
action_pad_modelast
action_expert_styleopenwam
action_mot_backbone_pretrained_path/cpfs/huangsq/VAM_Learn_from_Human_Video/src/vam/models/pretrained/ActionMoT_openwam_linear_interp_Wan22_alphascale_1024hdim.pt
mask_variantv2
mask_tail_padding_lossTrue
full_reference_videoTrue
max_ref_frames65
reference_dropout0.1
bridge_exclude_full_refTrue
extra_inputsvace_reference_image,action_trajectory
target_camerahead_camera
reference_camerahead_camera
resize_modepad
backboneti2v
model_paths["models/Wan-AI/Wan2.2-TI2V-5B/diffusion_pytorch_model-00001-of-00003.safetensors","models/Wan-AI/Wan2.2-TI2V-5B/diffusion_pytorch_model-00002-of-00003.safetensors","models/Wan-AI/Wan2.2-TI2V-5B/diffusion_pytorch_model-00003-of-00003.safetensors","models/Wan-AI/Wan2.2-TI2V-5B/models_t5_umt5-xxl-enc-bf16.pth","models/Wan-AI/Wan2.2-TI2V-5B/Wan2.2_VAE.pth"]
tokenizer_pathmodels/Wan-AI/Wan2.2-TI2V-5B/google/umt5-xxl
trainable_modelsdit
learning_rate5e-05
action_lrNone
weight_decay0.01
warmup_steps500
max_steps10000
num_epochs1
batch_size1
gradient_accumulation_steps1
dataset_repeat1
dataset_num_workers8
use_gradient_checkpointingTrue
save_steps1000
val_steps200
video_log_steps1000
max_val_samples20
lambda_video1
lambda_action1
video_dim3072
action_dit_dim1024
action_dit_ffn_dim4096
action_dit_num_heads24
action_dit_num_layers30
proprio_dropout0.3
window_stride1
val_ratio0.1
output_path/cpfs/huangsq/VAM_Learn_from_Human_Video/src/vam/models/train/wuji_hand_gesture_vam_ti2v5b_30L_openwam_fastwam_f33_r4_h32_ref65_refdrop10_mask_v2_propdrop30
wandb_projectwuji_hand_gesture
wandb_run_namewuji_hand_gesture_vam_ti2v5b_30L_mask_v2_openwam_refdrop_0.1_propdrop_0.3_0526_1545

Input/Output Contract

Expected inputs:

text
prompt:                 "the robot performs hand gesture {label}"
target camera:          head_camera
reference camera:       head_camera
target video frames:    33
full reference frames:  65
image resolution:       256 x 256
action horizon:         32
action/proprio dim:     20 / 20

Expected outputs:

text
robot-view target video rollout
20-D absolute robot action targets

Masking Note

This run uses mask_variant=v2 with the sequence layout:

text
[ref_video | first_frame | gen_video | action]

The run also records bridge_exclude_full_ref=True for provenance. For this OpenWAM/ActionMoT path, the active v2/v3 distinction is the mask_variant listed above.