datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
conceptual-captions-cc12m-llavanext
Dataset Card for conceptual-captions-cc12m-llavanext
Dataset Summary
This is a data of 21,930,344 synthetic captions for 10,965,172 images from conceptual_12m. In the interest of reproducibility, an archive found here on Huggingface was used (cc12m-wds). The captions were produced using llama3-llava-next-8b inferenced in float16, followed by cleanup and shortening with Meta-Llama-3-8B.
Languages
The captions are in English.
Data Instances
An… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/conceptual-captions-cc12m-llavanext.conceptual_captions_3m_zh_tiny_0
Dataset Card for "conceptual_captions_3m_zh_tiny_0"
More Information needed
conceptual_captions_3m_zh_tiny_4
Dataset Card for "conceptual_captions_3m_zh_tiny_4"
More Information needed
conceptual_captions_3m_zh_tiny_2
Dataset Card for "conceptual_captions_3m_zh_tiny_2"
More Information needed
conceptual_captions_3m_zh_tiny_5
Dataset Card for "conceptual_captions_3m_zh_tiny_5"
More Information needed
conceptual_captions_3m_zh_tiny_1
Dataset Card for "conceptual_captions_3m_zh_tiny_1"
More Information needed
conceptual_captions_3m_zh_tiny_3
Dataset Card for "conceptual_captions_3m_zh_tiny_3"
More Information needed
conceptual_captions_3m_zh_tiny_6
Dataset Card for "conceptual_captions_3m_zh_tiny_6"
More Information needed
