Camellia054/ShareGPT4V
News [2024/5/8] We released ShareGPT4Video, a large-scale video-caption dataset, with 40K captions annotated by GPT4V and 4.8M captions annotated by our ShareCaptioner-Video. The total videos last with 300 hours and 3000 hours separately! ShareGPT4V 1.2M Dataset Card Dataset details Dataset type: ShareGPT4V Captions 1.2M is a set of GPT4-Vision-powered multi-modal captions data. It is constructed to enhance modality alignment and fine-grained visual… See the full description on the dataset page: https://huggingface.co/datasets/Camellia054/ShareGPT4V.
News
[2024/5/8] We released [ShareGPT4Video](https://sharegpt4video.github.io/), a large-scale video-caption dataset, with 40K captions annotated by GPT4V and 4.8M captions annotated by our ShareCaptioner-Video. The total videos last with 300 hours and 3000 hours separately!
ShareGPT4V 1.2M Dataset Card
Dataset details
Dataset type: ShareGPT4V Captions 1.2M is a set of GPT4-Vision-powered multi-modal captions data.
It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Multi-Modal Models (LMMs) during both the pre-training and supervised fine-tuning stages. This advancement aims to bring LMMs towards GPT4-Vision capabilities.
- sharegpt4vinstructgpt4-vision_cap100k.json is generated by GPT4-Vision (ShareGPT4V).
- share-captionercocolcssam1246k_1107.json is generated by our Share-Captioner trained on GPT4-Vision-generated data (ShareGPT4V-PT).
- sharegpt4vmix665kcap23kcoco-ap9klcs3ksam9kdiv2k.json is curated from sharegpt4vinstructgpt4-vision_cap100k.json for the supervised fine-tuning stage.
Dataset date: ShareGPT4V Captions 1.2M was collected in 11.07 2023.
Paper or resources for more information: [Project] [Paper] [Code]
License: Attribution-NonCommercial 4.0 International It should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use
Intended use
Primary intended uses: The primary use of ShareGPT4V Captions 1.2M is research on large multimodal models and chatbots.
Primary intended users: The primary intended users of this dataset are researchers and hobbyists in computer vision, natural language processing, machine learning, and artificial intelligence.
