datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wds_imagenet1kwds_vtab-pcamMTADataset
Dataset Summary
MTADataset is a large-scale dataset designed for image inpainting.
For each image, we first employ Grounded-SAM to extract labels, bounding boxes, and masks.
Subsequently, we utilize LLaVA to provide detailed descriptions for approximately 5 masks per image, including information about their content and style.
For more details, please refer to our paper: 🌟 MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting
How to use
import os… See the full description on the dataset page: https://huggingface.co/datasets/huangjun12/MTADataset.wds_renderedsst2Backup_MTmtp
