CoolFace
20 results

recap

RLinf /RECAP-Libero10-Task0-48succ-Data2 likes17k downloads5mo agoHugging Facelmms-lab /LLaVA-ReCap-CC12Mimage1M<n<10M9 likes16k downloads2y agoHugging FaceUCSC-VLAA /Recap-DataComp-1B Dataset Card for Recap-DataComp-1B Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced LLaVA-1.5-LLaMA3-8B model to enhance the alignment and detail of textual descriptions. Dataset Details Dataset Description Our paper aims to bridge this community effort, leveraging the powerful and open-sourced LLaMA-3, a GPT-4 level LLM. Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/Recap-DataComp-1B.imagezero-shot-classification1B<n<10B205 likes9.9k downloads2y agoHugging FaceHon-Wong /VoRA-Recap-GLDv2-1.4Mtext1M<n<10M2 likes5.2k downloads1y agoHugging FaceDVijayan /vilhyra-recapture-000 likes4.5k downloads2h agoHugging Faceomni-research /Tarsier2-Recap-585Kgated Dataset Card for Tarsier2-Recap-585K Introduction ✨Tarsier2-Recap-585K✨ consists of 585K distinct video clips, lasting for 1972 hours in total, from open-source datasets (e.g. VATEX, TGIF, LSMDC, etc.) and each one with a detailed video description annotated by Tarsier2-7B, which beats GPT-4o in generating detailed and accurate video descriptions for video clips of 5~20 seconds (See the DREAM-1K Leaderboard). Experiments demonstrate its effectiveness in enhancing the… See the full description on the dataset page: https://huggingface.co/datasets/omni-research/Tarsier2-Recap-585K.videovideo-text-to-text22 likes3.6k downloads2y agoHugging Face