leinms/flickr30k-qwen3vl-baseline
Flickr30k Qwen3-VL Baseline Captions (Test Split) This dataset is based on the Mozilla/flickr30k-transformed-captions-gpt4o test split and contains 1,000 images from the original Flickr30k dataset.It includes both the original metadata and newly generated baseline captions produced using the Qwen3-VL-2B-Instruct vision-language model. Contents Each entry includes: image — the original Flickr30k image alt_text — GPT-4o transformed caption from Mozilla's… See the full description on the dataset page: https://huggingface.co/datasets/leinms/flickr30k-qwen3vl-baseline.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face