SEACrowd/bloom_captioning
This is a Bloom Library dataset developed for the image captioning task. It covers 74 languages indigenous to SEA overall, amounting to total data of 21K. This dataset belongs to a CC license, where its datapoints has specific license attached to it. Before using this dataloader, please accept the acknowledgement at https://huggingface.co/datasets/sil-ai/bloom-captioning and use huggingface-cli login for authentication. Languages abc, ahk, bfn, bjn, bkx, brb, brv, bya, bzi, ceb… See the full description on the dataset page: https://huggingface.co/datasets/SEACrowd/bloom_captioning.
This repository belongs to SEACrowd on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
