HuggingFaceM4/OBELICS
Dataset Card for OBELICS OBELICS is an open, massive, and curated collection of interleaved image-text web documents, containing 141M English documents, 115B text tokens, and 353M images, extracted from Common Crawl dumps between February 2020 and February 2023. The collection and filtering steps are described in our paper. Interleaved image-text web documents are a succession of text paragraphs interleaved by images, such as web pages that contain images. Models trained on… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/OBELICS.
This repository belongs to HuggingFaceM4 on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
