datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
augmented-recap-datacomp-3mThis is an experimental augmentation of about 3 million synthetic captions from Recap-Datacomp-1B. This dataset includes about 2 million multilingual captions.
It attempts to balance for gender stereotypes, added occupations, race, union membership, and religion to a subsample. We have also performed hair color and eye color balancing. It also includes some permutations of sentence orders, and modificaitons of the number of items ("Two" is changed to "Three", "Four", etc.)
We have also run… See the full description on the dataset page: https://huggingface.co/datasets/laion/augmented-recap-datacomp-3m.DataComp-1B_hudata-compare
