microsoft/git-large-coco
1063.1k
1---2language: en3license: mit4tags:5- vision6- image-captioning7model_name: microsoft/git-large-coco8pipeline_tag: image-to-text9---10 11# GIT (GenerativeImage2Text), large-sized, fine-tuned on COCO12 13GIT (short for GenerativeImage2Text) model, large-sized version, fine-tuned on COCO. It was introduced in the paper [GIT: A Generative Image-to-text Transformer for Vision and Language](https://arxiv.org/abs/2205.14100) by Wang et al. and first released in [this repository](https://github.com/microsoft/GenerativeImage2Text).14 15Disclaimer: The team releasing GIT did not write a model card for this model so this model card has been written by the Hugging Face team.16 17## Model description18 19GIT is a Transformer decoder conditioned on both CLIP image tokens and text tokens. The model is trained using "teacher forcing" on a lot of (image, text) pairs.20 21The goal for the model is simply to predict the next text token, giving the image tokens and previous text tokens.22 23The model has full access to (i.e. a bidirectional attention mask is used for) the image patch tokens, but only has access to the previous text tokens (i.e. a causal attention mask is used for the text tokens) when predicting the next text token.24 2526 27This allows the model to be used for tasks like:28 29- image and video captioning30- visual question answering (VQA) on images and videos31- even image classification (by simply conditioning the model on the image and asking it to generate a class for it in text).32 33## Intended uses & limitations34 35You can use the raw model for image captioning. See the [model hub](https://huggingface.co/models?search=microsoft/git) to look for36fine-tuned versions on a task that interests you.37 38### How to use39 40For code examples, we refer to the [documentation](https://huggingface.co/docs/transformers/main/model_doc/git#transformers.GitForCausalLM.forward.example).41 42## Training data43 44From the paper:45 46> We collect 0.8B image-text pairs for pre-training, which include COCO (Lin et al., 2014), Conceptual Captions47(CC3M) (Sharma et al., 2018), SBU (Ordonez et al., 2011), Visual Genome (VG) (Krishna et al., 2016),48Conceptual Captions (CC12M) (Changpinyo et al., 2021), ALT200M (Hu et al., 2021a), and an extra 0.6B49data following a similar collection procedure in Hu et al. (2021a).50 51=> however this is for the model referred to as "GIT" in the paper, which is not open-sourced.52 53This checkpoint is "GIT-large", which is a smaller variant of GIT trained on 20 million image-text pairs.54 55Next, the model was fine-tuned on COCO.56 57See table 11 in the [paper](https://arxiv.org/abs/2205.14100) for more details.58 59### Preprocessing60 61We refer to the original repo regarding details for preprocessing during training.62 63During validation, one resizes the shorter edge of each image, after which center cropping is performed to a fixed-size resolution. Next, frames are normalized across the RGB channels with the ImageNet mean and standard deviation.64 65## Evaluation results66 67For evaluation results, we refer readers to the [paper](https://arxiv.org/abs/2205.14100).