phi-9/re-recap
re-recap re-recap is an open vision-language dataset pairing live web image URLs with dual captions: the original baseline recaption from Recap-DataComp-1B and an expanded, highly detailed visual description generated directly from image pixels using Qwen3.8-27B. The dataset contains 9,717,471 verified image-text pairs partitioned into two distinct subsets based on caption length and prompt structure. Overview and Dataset Subsets To serve different modeling… See the full description on the dataset page: https://huggingface.co/datasets/phi-9/re-recap.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face