datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laion_text_debiased_60MFilter zxbsmk/laion_text_debiased_60M by image size and get 512 subset(12,009,641 pairs), 768 subset(4,915,850 pairs), 1024 subset(1,985,026 pairs).
Recap-Long-Laion
Dataset Card for Recap-Long-Laion
Dataset Description
This dataset consists of long captions of ~49M images from LAION-5B dataset. The long captions are generated by pre-trained Multi-modality Large Language Models (ShareGPT4V/InstructBLIP/LLava1.5) with the text prompt "Describe the image in detail".
Licensing Information
We distribute the image url with long captions under a standard Creative Common CC-BY-4.0 license. The individual images are under their own… See the full description on the dataset page: https://huggingface.co/datasets/weiwu-ww/Recap-Long-Laion.Laion_aesthetics_5plus_1024_33M_csvlaion-occupation
LAION Occupation
This dataset is a subset of LAION-2B-en containing 1.8M samples, each assigned to one of 153 occupations. This dataset was curated as part of our investigation into gender-occupation biases in LAION presented in Fair Diffusion.
For downloading the images, check out img2dataset.
Data Collection
We identified relevant images in the dataset by computing their CLIP similarity to a textual description of the target occupation. All descriptions were in the… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/laion-occupation.
