CoolFace
Datasetpublic

LamTNguyen/ridgelora-stage2-imposebase-train160-50k-20260824

Stage-2 ControlNet retraining with the frozen IMPOSE base This experiment retrains only Stage 2 for RidgeLoRA-FP. Stage 1 is the IMPOSE checkpoint and is not retrained. The run started on 2026-08-24 on TPU VM t1v-n-d3df3356-w-0 (TPU v5p-8, four XLA devices). An initial Stage-1-from-scratch job was stopped at step 575 after correcting the scope. It produced no scheduled checkpoint and is not used in any result; its log is retained only as an audit trail. Frozen IMPOSE… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-stage2-imposebase-train160-50k-20260824.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes88downloads
Dataset Card

Stage-2 ControlNet retraining with the frozen IMPOSE base

This experiment retrains only Stage 2 for RidgeLoRA-FP. Stage 1 is the IMPOSE checkpoint and is not retrained. The run started on 2026-08-24 on TPU VM t1v-n-d3df3356-w-0 (TPU v5p-8, four XLA devices).

An initial Stage-1-from-scratch job was stopped at step 575 after correcting the scope. It produced no scheduled checkpoint and is not used in any result; its log is retained only as an audit trail.

Frozen IMPOSE components

  • —checkpoint: /home/thanhlamtba31/FingenBench/IMPOSE-official/models/fingerprint_c2cl_512/model_unsw_flat.ckpt
  • —SHA-256: 3048e170b8fd4687a18b54ca823dcb01ca1a0f5fe2f5affc297a1f00072763ab
  • —frozen VQ-VAE: 13,855,580 parameters
  • —frozen IMPOSE/base U-Net: 16,334,339 parameters
  • —trainable Stage-2 ControlNet: 7,039,760 parameters

A fresh ControlNet copies 116 compatible U-Net tensors (6,676,736 parameters). Its 32 zero-convolution and hint-path tensors (363,024 parameters) retain their released initialization. The trained ControlNet tensors in the official cross-modal checkpoint are not loaded.

Subject-isolated Stage-2 data

  • —split SHA-256: 063d295927798202f95f8c5ed0b855c102971bed3ec26b0792b1c5602d61ba65
  • —split: 160 train / 20 validation / 20 test subjects
  • —training sensors: A, B, and F
  • —training image/ridge pairs: 4,722
  • —validation/test subjects used in Stage-2 training: 0/0
  • —latent manifest SHA-256: acd5738a95cc97c189ad74468777e0945b06bcf728ffd1fa0d86446b751dc9cf

The frozen IMPOSE VQ-VAE encodes targets to 3 x 128 x 128. Paired 1 x 512 x 512 ridge conditions use Sauvola thresholding with window 11, k=0.007, R=128, followed by a 2x2 opening on inverted ridge foreground.

Active Stage-2 run

  • —output: /home/thanhlamtba31/FingenBench/experiments/clean_retrain/stage2_train_imposebase_train160_50k
  • —target: 50,000 steps
  • —global batch: 256 over four TPU devices
  • —loss: L1 epsilon prediction
  • —diffusion: 1,000 linear steps, beta 0.0015 to 0.0155
  • —LR: 5e-5 with 500-step warmup
  • —seed: 20260608
  • —EMA: 0.9999
  • —gradient clipping: 1.0
  • —checkpoints: every 10,000 steps
  • —initial smoke throughput: approximately 2.67 steps/s

The launcher is train_stage2_xla.py. It uses BF16 model computation with FP32 trainable weights, EMA, and Adam moments. The frozen IMPOSE U-Net and VQ-VAE are never updated.

Completion and Hugging Face upload

After step 50,000, package and SHA-256 verify the following before upload:

  • —final Stage-2 checkpoint and optimizer/EMA state
  • —train_metrics.jsonl and the final run report
  • —split, latent-manifest, and condition audit reports
  • —train_stage2_xla.py and retrain_impose_xla.py
  • —this README and exact launch/resume commands

Upload to a new public dataset repo under LamTNguyen. The upload script must read authentication only from the HF_TOKEN environment variable; no token is stored in source, README, shell history, or logs. After upload, compare local SHA-256 values with the files downloaded from the Hub before declaring the archive complete.

Next actions

  1. 1.Generate QC100 using the 50K Stage-2 EMA and frozen IMPOSE base.
  2. 2.If QC passes, generate 5,000 matched-budget images and run the recognizer.
  3. 3.Run direct cross-sensor synthetic downstream evaluation.
  4. 4.Run causal ridge-preprocessor, latent-resolution, and LoRA placement/rank ablations.