Lo/adapt-pre-trained-VL-models-to-text-data-Wikipedia-finetune
The Wikipedia finetune data used to train visual features for the adaption of vision-and-language models to text-only tasks in the paper "How to Adapt Pre-trained Vision-and-Language Models to a Text-only Input?". The data has been created from the "20200501.en" revision of the wikipedia dataset on Huggingface.
023
