mahadev-balla/isl-conditional-diffusion-linear-128
01
ISL Conditional Diffusion — Linear Noise Schedule
A class-conditioned DDPM trained from scratch to generate 128×128 RGB images of Indian Sign Language hand gestures using a linear noise schedule.
The model uses class conditioning on one of 35 ISL classes and supports classifier-free guidance during sampling.
The complete implementation and experiments are available in the GitHub repository.
Model
- DDPM with UNet2D architecture
- 128×128 resolution
- RGB images
- Class-conditional generation
- 35 ISL classes
- Classifier-free guidance (CFG)
- EMA weights
Training
- Dataset: 42,000 images (1,200 images per class × 35 classes)
- Noise schedule: linear
- Batch size: 64
- Learning rate: 1e-4
- Mixed precision: fp16
- Training steps: 65,000
- EMA decay: 0.9999
- CFG label dropout: 0.15
- Data augmentation: enabled
Sampling
- Default sampler: DDIM
- Training diffusion timesteps: 1,000
- Routine inference steps: 100
- Evaluation inference steps: 50
- Default guidance scale: 1.0
- Random seed: 42
Results
At the FID-optimal guidance scale of 1.0:
- FID: 58.25
- Semantic accuracy: 98.0%
The model provides strong class control under the reported evaluation setting.
