AterMors/Swin2-GTP2_art-caption
021
Image Captioning Model created with VisionEncoderDecoderModel architecture using "microsoft/swinv2-base-patch4-window12to16-192to256-22kto1k-ft" as imageencoder and "openai/gpt2" as textdecoder. It has been trained on a variant of the WikiArt dataset that can be found at "AterMors/wikiart_recaption".
