datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Recap-DataComp-1B
Dataset Card for Recap-DataComp-1B
Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced LLaVA-1.5-LLaMA3-8B model to enhance the alignment and detail of textual descriptions.
Dataset Details
Dataset Description
Our paper aims to bridge this community effort, leveraging the powerful and open-sourced LLaMA-3, a GPT-4 level LLM.
Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/Recap-DataComp-1B.recap-datacomp-12m-wdsdatacomp_recap_metadata2Recap-Datacomp-1B_tars_part7Final size: 7,236,721, samples per tar: 10000
Recap-DataComp-1B_split_3Recap-DataComp-1B_split_4Recap-DataComp-1B-FoodOrDrink
Recap-DataComp-1B: Food or Drink
A filtered subset of Recap-DataComp-1B containing 106,230,157 rows classified as food/drink content, enriched with structured food/drink extraction from FoodExtract-v2.
Overview
Count
Percentage
Total rows
106,230,157
100%
Food/drink (Stage 5 label)
96,618,895
91.0%
Not food/drink (Stage 5 label)
9,611,262
9.0%
FoodExtract (re_caption): food/drink
79,519,489
74.9%
FoodExtract (re_caption): not food/drink
26,710,156… See the full description on the dataset page: https://huggingface.co/datasets/mrdbourke/Recap-DataComp-1B-FoodOrDrink.Recap-Datacomp-1B_tars_part9Final size: 7,240,328, samples per tar: 10000
Recap-DataComp-1B_split_5Recap-Datacomp-1B_tars_part8Final size: 7,250,604, samples per tar: 10000
Recap-Datacomp-1B_tars_part13Recap-DataComp-1B_split_7Recap-Datacomp-1B_tars_part14Recap-DataComp-1B_split_8Recap-DataComp-1B_split_2Recap-DataComp-1B_split_1Recap-DataComp-1B_split_6Recap-DataComp-100K
Description
Recap-DataComp-100K is a subset of UCSC-VLAA/Recap-DataComp-1B.
This dataset aims to ease the development of vision-language models by providing a readily-available small collection of image-text pairs.
Use this dataset for sanity checks, developing POCs, or other quick multimodal dev. For serious model training please refer to the original repo linked above.
Citation
Always cite the original authors . I've copied their citation info here for your… See the full description on the dataset page: https://huggingface.co/datasets/nnethercott/Recap-DataComp-100K.datacomp-recap-qwen3p5-35b-a3b
DataComp recaptions with Qwen3.5-35B-A3B
Dataset datacomp: 325.623 Million caption rows.
This public caption-only repository contains 325,623,447 generated caption rows for 325,472,073 normalized-URL identities. It unifies the selected existing recap pack, the later generated/recovery records, and the managed legacy v2 recap pack without exposing historical surface names. Distinct stripped captions for the same URL are retained; only exact normalized-URL plus stripped-caption… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/datacomp-recap-qwen3p5-35b-a3b.recap-datacomp-384-1Mrecap-datacomp-1B-moondream-test1
