CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01adams-story /datacomp200m Datacomp200m This is a smaller version of the datacomp_1b dataset. Filtering was done by taking all rows that had self similarity (inner product) above 0.32. This resulted in 213009083 (213 million) rows. The results of the datacomp paper suggest that filtering by CLIP score is better than random sampling. Included in this repo are search indices created using autofaiss, over the text and image embeddings. There are two ways to access metadata, either in .parquet files in the… See the full description on the dataset page: https://huggingface.co/datasets/adams-story/datacomp200m.image100M<n<1B3 likes110k downloads3y agoHugging Face02mlfoundations /datacomp_xlarge DataComp XLarge Pool This repository contains metadata files for the xlarge pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_xlarge.image10B<n<100B21 likes39k downloads3y agoHugging Face03mlfoundations /datacomp_1b DataComp-1B This repository contains metadata files for DataComp-1B. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_1b.image1B<n<10B53 likes15k downloads3y agoHugging Face04UCSC-VLAA /Recap-DataComp-1B Dataset Card for Recap-DataComp-1B Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced LLaVA-1.5-LLaMA3-8B model to enhance the alignment and detail of textual descriptions. Dataset Details Dataset Description Our paper aims to bridge this community effort, leveraging the powerful and open-sourced LLaMA-3, a GPT-4 level LLM. Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/Recap-DataComp-1B.imagezero-shot-classification1B<n<10B205 likes9.5k downloads2y agoHugging Face05mlfoundations /datacomp_large DataComp Large Pool This repository contains metadata files for the large pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_large.image1B<n<10B6 likes6.2k downloads3y agoHugging Face06cornuHGF /recap-datacomp-12m-wdsimage10M<n<100M0 likes2.3k downloads1y agoHugging Face07mlfoundations /datacomp_small DataComp Small Pool This repository contains metadata files for the small pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_small.image10M<n<100M6 likes1.7k downloads3y agoHugging Face08mlfoundations /datacomp_medium DataComp Medium Pool This repository contains metadata files for the medium pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_medium.image100M<n<1B3 likes1.2k downloads3y agoHugging Face09nielsr /datacomp-small-with-text-embeddings Dataset Card for "datacomp-small-with-text-embeddings" More Information needed image10M<n<100M0 likes1.1k downloads3y agoHugging Face10thaottn /datacomp-medium-pool-translatedimage100M<n<1B0 likes1k downloads1y agoHugging Face11Leonardo6 /Datacomp-10m-embeddingimage1M<n<10M0 likes908 downloads2y agoHugging Face12laion /datacomp-hqimage100M<n<1B14 likes677 downloads3y agoHugging Face13nielsr /datacomp-small-with-embeddings Dataset Card for "datacomp-small-with-embeddings" More Information needed image10M<n<100M0 likes654 downloads3y agoHugging Face14nielsr /datacomp-small-with-embeddings-and-cluster-labels Dataset Card for "datacomp-small-with-embeddings-and-cluster-labels" More Information needed image10M<n<100M0 likes533 downloads3y agoHugging Face15Leonardo6 /Datacomp-10m-embedimage10M<n<100M0 likes479 downloads2y agoHugging Face16nielsr /datacomp-small-filtered Dataset Card for "datacomp-small-filtered" This is the DataComp-small dataset with CLIP-large-patch14 image embeddings added, as well as: captions filtered for English using a FastText model captions filtered to have at least complexity of 1 image1M<n<10M1 likes383 downloads3y agoHugging Face17umd-vt-nyu /datacomp_recap_metadata2image100M<n<1B2 likes338 downloads2y agoHugging Face18fondant-ai /datacomp-small-clip Production-ready data processing made easy and shareable Explore the Fondant docs » Dataset Card for fondant-ai/datacomp-small-clip This is a dataset containing image urls and their CLIP embeddings, based on the datacomp_small dataset, and processed with fondant. Dataset Details Dataset Description Large (image) datasets are often unwieldy to use due to their… See the full description on the dataset page: https://huggingface.co/datasets/fondant-ai/datacomp-small-clip.imageimage-to-text10M<n<100M14 likes322 downloads3y agoHugging Face19Hennara /Recap-DataComp-1B_split_3image100M<n<1B0 likes319 downloads2y agoHugging Face20Hennara /Recap-DataComp-1B_split_4image100M<n<1B0 likes300 downloads2y agoHugging Face21cornuHGF /datacomp-smallimage10M<n<100M1 likes298 downloads1y agoHugging Face22minhnguyent546 /datacomp_large_vie_imagesThis repository contains images downloaded with img2dataset for minhnguyent546/datacomp_large_vie_filtered2. image1M<n<10M0 likes275 downloads6mo agoHugging Face23mrdbourke /Recap-DataComp-1B-FoodOrDrink Recap-DataComp-1B: Food or Drink A filtered subset of Recap-DataComp-1B containing 106,230,157 rows classified as food/drink content, enriched with structured food/drink extraction from FoodExtract-v2. Overview Count Percentage Total rows 106,230,157 100% Food/drink (Stage 5 label) 96,618,895 91.0% Not food/drink (Stage 5 label) 9,611,262 9.0% FoodExtract (re_caption): food/drink 79,519,489 74.9% FoodExtract (re_caption): not food/drink 26,710,156… See the full description on the dataset page: https://huggingface.co/datasets/mrdbourke/Recap-DataComp-1B-FoodOrDrink.imagetext-classification100M<n<1B1 likes248 downloads6mo agoHugging Face24partitionsofunity /DataComp-1Bimage10K<n<100K0 likes234 downloads2y agoHugging Face25apehex /ascii-art-datacompdr-12m ASCII Art DataCompDR-12M Description This is a text-to-image dataset, where the images are actually ASCII art. The images and captions were sampled from DataCompDR-12M. The conversion was performed with the tool ascii-image-converter. Metadata homepage: https://github.com/apehex/scrapscii version: 0.1.0 Config Split Size Samples 'default' 'train' 4.1 GB 643072 'default' 'fixed' 552 MB 262144 The ASCII art in "fixed" all have a width of 64… See the full description on the dataset page: https://huggingface.co/datasets/apehex/ascii-art-datacompdr-12m.text100K<n<1M0 likes215 downloads1y agoHugging Face26jordangong /datacomp12m_allimage10M<n<100M1 likes213 downloads2y agoHugging Face27Hennara /Recap-DataComp-1B_split_5image100M<n<1B0 likes196 downloads2y agoHugging Face28Hennara /Recap-DataComp-1B_split_7image100M<n<1B0 likes168 downloads2y agoHugging Face29Hennara /Recap-DataComp-1B_split_8image100M<n<1B0 likes152 downloads2y agoHugging Face30nielsr /datacomp-small-with-embeddings-ca-filtered Dataset Card for "datacomp-small-with-embeddings-ca-filtered" This is the datacomp-small dataset, with CLIP-large-patch14 image embeddings added, as well as CA filtering: minimum caption complexity of 1 minimum 1 action in the caption image1M<n<10M0 likes151 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.