datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
datacomp200m
Datacomp200m
This is a smaller version of the datacomp_1b dataset.
Filtering was done by taking all rows that had self similarity (inner product) above 0.32. This resulted in 213009083 (213 million) rows.
The results of the datacomp paper suggest that filtering by CLIP score is better than random sampling.
Included in this repo are search indices created using autofaiss, over the text and image embeddings. There are two ways to access metadata, either in .parquet files in the… See the full description on the dataset page: https://huggingface.co/datasets/adams-story/datacomp200m.datacomp_pools
DataComp Pools
This repository contains metadata files for DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_pools.datacomp_xlarge
DataComp XLarge Pool
This repository contains metadata files for the xlarge pool of DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_xlarge.DataCompDR-1B
Dataset Card for DataCompDR-1B
This dataset contains synthetic captions, embeddings, and metadata for DataCompDR-1B.
The metadata has been generated using pretrained image-text models on DataComp-1B.
For details on how to use the metadata, please visit our github repository.
Dataset Details
Dataset Description
DataCompDR is an image-text dataset and an enhancement to the DataComp dataset.
We reinforce the DataComp dataset using our multi-modal… See the full description on the dataset page: https://huggingface.co/datasets/apple/DataCompDR-1B.Recap-DataComp-1B
Dataset Card for Recap-DataComp-1B
Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced LLaVA-1.5-LLaMA3-8B model to enhance the alignment and detail of textual descriptions.
Dataset Details
Dataset Description
Our paper aims to bridge this community effort, leveraging the powerful and open-sourced LLaMA-3, a GPT-4 level LLM.
Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/Recap-DataComp-1B.DataCompDR-12M
Dataset Card for DataCompDR-12M
This dataset contains synthetic captions, embeddings, and metadata for DataCompDR-12M.
The metadata has been generated using pretrained image-text models on a 12M subset of DataComp-1B.
For details on how to use the metadata, please visit our github repository.
The dataset with the original captions is now available at mlfoundations/DataComp-12M.
The UIDs per shards match between mlfoundations/DataComp-12M and apple/DataCompDR-12M.… See the full description on the dataset page: https://huggingface.co/datasets/apple/DataCompDR-12M.datacomp_large
DataComp Large Pool
This repository contains metadata files for the large pool of DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_large.datacomp_1b
DataComp-1B
This repository contains metadata files for DataComp-1B. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_1b.DataComp-12M
Dataset Card for DataComp-12M
This dataset contains a 12M subset of DataComp-1B-BestPool.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Image-text models trained on DataComp-12M are significantly better than on CC-12M/YFCC-15M as well as DataComp-Small/Medium.
DataComp-12M was introduced in MobileCLIP paper and along with the reinforced dataset DataCompDR-12M.
The UIDs… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/DataComp-12M.TiC-DataComp
Dataset Card for TiC-DataComp
This dataset containts metadata for TiC-DataComp benchmark for time-continual learning of image-text models.
The dataset containts timestamp information for DataComp-1B in the form of UIDs groupings by year/month sourced from the original CommonCrawl.
We also release UIDs for our TiC-DataCompNet and TiC-DataComp-Retrieval evaluations for continual learning of CLIP models.
For details on how to use the metadata, please visit our github repository.… See the full description on the dataset page: https://huggingface.co/datasets/apple/TiC-DataComp.datacomp_small
DataComp Small Pool
This repository contains metadata files for the small pool of DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_small.recap-datacomp-12m-wdsdatacomp_medium
DataComp Medium Pool
This repository contains metadata files for the medium pool of DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_medium.DataCompDR-12M-bf16
Dataset Card for DataCompDR-12M-BFloat16
This dataset contains synthetic captions, embeddings, and metadata for DataCompDR-12M.
The metadata has been generated using pretrained image-text models on a 12M subset of DataComp-1B.
For details on how to use the metadata, please visit our github repository.
The dataset with the original captions is now available at mlfoundations/DataComp-12M.
The UIDs per shards match between mlfoundations/DataComp-12M and apple/DataCompDR-12M-bf16.… See the full description on the dataset page: https://huggingface.co/datasets/apple/DataCompDR-12M-bf16.datacomp-medium-pool-translatedDatacomp-10m-embeddingdatacomp_subdatacomp-small-with-text-embeddings
Dataset Card for "datacomp-small-with-text-embeddings"
More Information needed
datacomp-10m-embed-7bdatacomp-hqDataComp-12M-Images-256
DataComp-12M Translated
Description
Translated DataComp-12M. English captions were machine-translated to Russian using GigaChat3-10B-A1.8B.
Statistics
Metric
Count
Original dataset
12,561,027
Successfully downloaded
8,744,177 (69.6%)
Failed to download
3,497,631 (27.8%)
Failed to resize
319,219 (2.5%)
Full Dataset
*.parquet - Full metadata (1,257 files, 5 corrupted)
*.tar - Images resized to 256px
Corrupted Parquet… See the full description on the dataset page: https://huggingface.co/datasets/Alexator26/DataComp-12M-Images-256.datacomp-small-with-embeddings
Dataset Card for "datacomp-small-with-embeddings"
More Information needed
imagenet-1k-random-0.0-frac-1over4imagenet-1k-random20.0datacomp-small-with-embeddings-and-cluster-labels
Dataset Card for "datacomp-small-with-embeddings-and-cluster-labels"
More Information needed
imagenet-1k-random-30.0-frac-1over8Datacomp-10m-embeddatacomp-smallprovision_datacomp_imagesdatacomp-medium-12m
