datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flickr1k-sd-attn
Flickr1k VQA Dataset with Per-Word Attention Maps
This dataset is a 1K subset of the Flickr30k VQA dataset, including per-word attention maps extracted with Stable Diffusion v1.4.
📂 Structure
image: Main image associated with the question, stored in the images/ subfolder.
question: VQA question derived from Flickr30k captions.
answers: Ground-truth answers.
attention_images: A dictionary of per-word saliency images stored in the attention_images/ subfolder.
All images… See the full description on the dataset page: https://huggingface.co/datasets/lxasqjc/flickr1k-sd-attn.libero-attn-compare-task1-ep1-frame40last05-libero-attn-vis-compare-task0-ep1-frame50CLIP-Cross-Attn-MUX-Training-Data-Pack
CLIP-MUX Training Data Pack
This repository is the data/metadata companion for reproducing the CLIP-MUX / x-attention CLIP training pipeline.Used to train model: zer0int/CLIP-ViT-L-14-Cross-Attn-Read-NoRead-ModeMUX
This is not one monolithic dataset under one license. Each component is independently scoped and carries its own LICENSE or NOTICE file.
The pack intentionally separates:
assets that can be redistributed directly;
runtime labels/manifests derived from upstream… See the full description on the dataset page: https://huggingface.co/datasets/zer0int/CLIP-Cross-Attn-MUX-Training-Data-Pack.flickr30k_attn_ft
