arikatokachi/indic-multilingual-image-captions
Indic Multilingual Image Caption Dataset This dataset contains 4,500 unique images with captions in: English Hindi Bengali Tamil Source composition 3,000 images from COCO Caption 2017 1,500 images from TextCaps Each image is stored once and paired with four multilingual caption variants. Hindi, Bengali and Tamil captions were generated from the selected English captions using the NLLB-200 distilled translation model. Intended use The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/arikatokachi/indic-multilingual-image-captions.
This repository belongs to arikatokachi on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
