CoolFace
Datasetpublic

wikimedia/wit_base

Dataset Card for WIT Dataset Summary Wikimedia's version of the Wikipedia-based Image Text (WIT) Dataset, a large multimodal multilingual dataset. From the official blog post: The core training data is taken from the Wikipedia Image-Text (WIT) Dataset, a large curated set of more than 37 million image-text associations extracted from Wikipedia articles in 108 languages that was recently released by Google Research. The WIT dataset offers extremely valuable data… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/wit_base.

sourceHugging Facecc-by-sa-4.0updated 4y agoView on Hugging Face
75likes15kdownloads
settings

This repository belongs to wikimedia on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namewit_base
visibilitypublic
licencecc-by-sa-4.0
gatedno
ownerwikimedia
Account settings
wikimedia/wit_base · CoolFace