huggingface/test-big-dataset
Dataset Card for Danish WIT Dataset Summary Google presented the Wikipedia Image Text (WIT) dataset in July 2021, a dataset which contains scraped images from Wikipedia along with their descriptions. WikiMedia released WIT-Base in September 2021, being a modified version of WIT where they have removed the images with empty "reference descriptions", as well as removing images where a person's face covers more than 10% of the image surface, along with inappropriate… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/test-big-dataset.
0362
Duplicate from severo/danish-wit
