datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
datacomp200m
Datacomp200m
This is a smaller version of the datacomp_1b dataset.
Filtering was done by taking all rows that had self similarity (inner product) above 0.32. This resulted in 213009083 (213 million) rows.
The results of the datacomp paper suggest that filtering by CLIP score is better than random sampling.
Included in this repo are search indices created using autofaiss, over the text and image embeddings. There are two ways to access metadata, either in .parquet files in the… See the full description on the dataset page: https://huggingface.co/datasets/adams-story/datacomp200m.datacomp_xlarge
DataComp XLarge Pool
This repository contains metadata files for the xlarge pool of DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_xlarge.Recap-DataComp-1B
Dataset Card for Recap-DataComp-1B
Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced LLaVA-1.5-LLaMA3-8B model to enhance the alignment and detail of textual descriptions.
Dataset Details
Dataset Description
Our paper aims to bridge this community effort, leveraging the powerful and open-sourced LLaMA-3, a GPT-4 level LLM.
Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/Recap-DataComp-1B.datacomp_large
DataComp Large Pool
This repository contains metadata files for the large pool of DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_large.datacomp_1b
DataComp-1B
This repository contains metadata files for DataComp-1B. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_1b.datacomp_small
DataComp Small Pool
This repository contains metadata files for the small pool of DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_small.recap-datacomp-12m-wdsdatacomp_medium
DataComp Medium Pool
This repository contains metadata files for the medium pool of DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_medium.datacomp-medium-pool-translateddatacomp-small-with-text-embeddings
Dataset Card for "datacomp-small-with-text-embeddings"
More Information needed
datacomp-hqdatacomp-small-with-embeddings
Dataset Card for "datacomp-small-with-embeddings"
More Information needed
datacomp-small-with-embeddings-and-cluster-labels
Dataset Card for "datacomp-small-with-embeddings-and-cluster-labels"
More Information needed
datacomp-smalldatacomp_recap_metadata2datacomp-small-filtered
Dataset Card for "datacomp-small-filtered"
This is the DataComp-small dataset with CLIP-large-patch14 image embeddings added, as well as:
captions filtered for English using a FastText model
captions filtered to have at least complexity of 1
Recap-DataComp-1B_split_3datacomp_large_vie_imagesThis repository contains images downloaded with img2dataset for minhnguyent546/datacomp_large_vie_filtered2.
DataComp-1Bdatacomp12m
DataComp-12M
This repository contains DataComp-12M metadata in the same format as DataComp-1B. The data was filtered from DataComp-1B using the UIDs provided in apple/DataComp-12M.
Recap-DataComp-1B_split_4Recap-DataComp-1B-FoodOrDrink
Recap-DataComp-1B: Food or Drink
A filtered subset of Recap-DataComp-1B containing 106,230,157 rows classified as food/drink content, enriched with structured food/drink extraction from FoodExtract-v2.
Overview
Count
Percentage
Total rows
106,230,157
100%
Food/drink (Stage 5 label)
96,618,895
91.0%
Not food/drink (Stage 5 label)
9,611,262
9.0%
FoodExtract (re_caption): food/drink
79,519,489
74.9%
FoodExtract (re_caption): not food/drink
26,710,156… See the full description on the dataset page: https://huggingface.co/datasets/mrdbourke/Recap-DataComp-1B-FoodOrDrink.Recap-DataComp-1B_split_2Recap-DataComp-1B_split_5datacomp12m_allRecap-DataComp-1B_split_7Recap-DataComp-1B_split_8small_datacomp_alldatacomp_chunked_128Recap-DataComp-1B_split_6
