CoolFace
Datasetpublic

HuggingFaceM4/OBELICS

Dataset Card for OBELICS OBELICS is an open, massive, and curated collection of interleaved image-text web documents, containing 141M English documents, 115B text tokens, and 353M images, extracted from Common Crawl dumps between February 2020 and February 2023. The collection and filtering steps are described in our paper. Interleaved image-text web documents are a succession of text paragraphs interleaved by images, such as web pages that contain images. Models trained on… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/OBELICS.

sourceHugging Facecc-by-4.0updated 3y agoView on Hugging Face
174likes17kdownloads

No commit history came back for main. The revision may not exist, or the source declined the request.