datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ramanv-image-captions-6multilingual-image-captions-text
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
multilingual_image_captions_text
This dataset provides text-only multilingual image annotations derived from the original 'multilingual-image-annotations' collection, with binary image data removed for efficient loading. It contains 464 rows of English and multilingual captions generated by the google/gemma-4-31B-it model across seven languages. The content covers diverse… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-image-captions-text.ImageCaptions-7M-Translations-Arabic-subset-150000
