simonko912/pikadiffusion
3215

Pikadiffusion
Model details
- Parameters: ~129M
- Resolution: 128x128
- Dataset: 6576 Pokémon images + captions
- Epochs: 100
- Text encoder: OpenAI CLIP ViT-B/32
- Diffusion: DDPM
- EMA: Yes
- Loss after training: 0.017578418961433816
Datasets
- hf.co/TeeA/Pokemon-Captioning-Classification
- hf.co/GabeHD/pokemon-type-captions
- github.com/rileynwong/pokemon-images-dataset-by-type
Notes
lowram.py is reccomended for devices under 4GB of ram. Uses around 2gb peak and 1gb while generating.<br> Check out the space! Takes 5s per image https://huggingface.co/spaces/simonko912/pikadiffusion <Gallery />
<br>
