wikimedia/wit_base
Dataset Card for WIT Dataset Summary Wikimedia's version of the Wikipedia-based Image Text (WIT) Dataset, a large multimodal multilingual dataset. From the official blog post: The core training data is taken from the Wikipedia Image-Text (WIT) Dataset, a large curated set of more than 37 million image-text associations extracted from Wikipedia articles in 108 languages that was recently released by Google Research. The WIT dataset offers extremely valuable data… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/wit_base.
This repository belongs to wikimedia on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
