laicsiifes/flickr8k-pt-br
🎉 Flickr8K Dataset Translation for Portuguese Image Captioning 💾 Dataset Summary Flickr8K Portuguese Translation, a multimodal dataset for Portuguese image captioning with 8,000 images, each accompanied by five descriptive captions that have been generated by human annotators for every individual image. The original English captions were rendered into Portuguese through the utilization of the Google Translator API. 🧑💻 Hot to Get Started with the… See the full description on the dataset page: https://huggingface.co/datasets/laicsiifes/flickr8k-pt-br.
🎉 Flickr8K Dataset Translation for Portuguese Image Captioning
💾 Dataset Summary
Flickr8K Portuguese Translation, a multimodal dataset for Portuguese image captioning with 8,000 images, each accompanied by five descriptive captions that have been generated by human annotators for every individual image. The original English captions were rendered into Portuguese through the utilization of the Google Translator API.
🧑💻 Hot to Get Started with the Dataset
from datasets import load_dataset
dataset = load_dataset('laicsiifes/flickr8k-pt-br')✍️ Languages
The images descriptions in the dataset are in Portuguese.
🧱 Dataset Structure
📝 Data Instances
An example looks like below:
{
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=500x399>,
'filename': 'ad5f9478d603fe97.jpg',
'caption': [
'Um cachorro preto está correndo atrás de um cachorro branco na neve.',
'Cachorro preto perseguindo cachorro marrom na neve',
'Dois cães perseguem um ao outro pelo chão nevado.',
'Dois cachorros brincam juntos na neve.',
'Dois cães correndo por um corpo de água baixo.'
]
}🗃️ Data Fields
The data instances have the following fields:
image: aPIL.Image.Imageobject containing image.filename: astrcontaining name of image file.caption: alistofstrcontaining 5 captions related to image.
✂️ Data Splits
The dataset is partitioned using the Karpathy splitting appoach for Image Captioning (Karpathy and Fei-Fei, 2015).
📋 BibTeX entry and citation info
@misc{bromonschenkel2024flickr8kpt,
title = {Flickr8K Dataset Translation for Portuguese Image Captioning},
author = {Bromonschenkel, Gabriel and Oliveira, Hil{\'a}rio and Paix{\~a}o, Thiago M.},
howpublished = {\url{https://huggingface.co/datasets/laicsiifes/flickr8k-pt-br}},
publisher = {Hugging Face},
year = {2024}
}