CoolFace
Datasetpublic

HuggingFaceM4/OBELICS

Dataset Card for OBELICS OBELICS is an open, massive, and curated collection of interleaved image-text web documents, containing 141M English documents, 115B text tokens, and 353M images, extracted from Common Crawl dumps between February 2020 and February 2023. The collection and filtering steps are described in our paper. Interleaved image-text web documents are a succession of text paragraphs interleaved by images, such as web pages that contain images. Models trained on… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/OBELICS.

sourceHugging Facecc-by-4.0updated 3y agoView on Hugging Face
174likes17kdownloads
settings

This repository belongs to HuggingFaceM4 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameOBELICS
visibilitypublic
licencecc-by-4.0
gatedno
ownerHuggingFaceM4
Account settings
HuggingFaceM4/OBELICS · CoolFace