CoolFace
Datasetpublic

linyq/laion_text_debiased_100M

100M Text Debiased Subset from LAION 2B Captions in LAION-2B have a significant bias towards describing visual text content embedded in the images. Released CLIP models have strong text spotting bias in almost every style of web images, resulting in the CLIP-filtering datasets inherently biased towards visual text dominant data. CLIP models easily learn text spotting capacity from parrot captions while failing to connect the vision-language semantics, just like a text spotting… See the full description on the dataset page: https://huggingface.co/datasets/linyq/laion_text_debiased_100M.

sourceHugging Facecc-by-4.0updated 3y agoView on Hugging Face
0likes121downloads
4 commits on main
a9ac1063y ago

Update README.md

linyq
aa8c91b3y ago

Upload dataset (part 00001-of-00002)

linyq
6b5052c3y ago

Upload dataset (part 00000-of-00002)

linyq
3212a3f3y ago

initial commit

linyq