CoolFace
Datasetpublic

Shanmh/ShareGPT4Video

ShareGPT4Video 4.8M Dataset Card Dataset details Dataset type: ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos. It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards the capabilities of GPT4V and Sora. sharegpt4video_40k.jsonl is generated by… See the full description on the dataset page: https://huggingface.co/datasets/Shanmh/ShareGPT4Video.

sourceHugging Facecc-by-nc-4.0updated 2mo agoView on Hugging Face
0likes17downloads
Dataset Card

ShareGPT4Video 4.8M Dataset Card

Dataset details

Dataset type: ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos.

It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards the capabilities of GPT4V and Sora.

  • —sharegpt4video_40k.jsonl is generated by GPT4-Vision (ShareGPT4Video).
  • —share-captioner-videomixkit-pexels-pixabay4814k_0417.json is generated by our ShareCaptioner-Video trained on GPT4-Vision-generated video-caption pairs.
  • —sharegpt4videomix181kvqa-153kshare-cap-28k.json is curated from sharegpt4videoinstructgpt4-visioncap40k.json for the supervised fine-tuning stage of LVLMs.
  • —llavav15mix665kwithvideochatgpt72k_share4video28k.json has replaced 28K detailed-caption-related data in VideoChatGPT with 28K high-quality captions from ShareGPT4Video. This file is utilized to validate the effectiveness of high-quality captions under the VideoLLaVA and LLaMA-VID models.

Dataset date:

ShareGPT4Video Captions 4.8M was collected in 4.17 2024.

Paper or resources for more information: [Project] [Paper] [Code] [ShareGPT4Video-8B]

License: Attribution-NonCommercial 4.0 International It should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use

Intended use

Primary intended uses: The primary use of ShareGPT4Video Captions 4.8M is research on large multimodal models and text-to-video models. Primary intended users: The primary intended users of this dataset are researchers and hobbyists in computer vision, natural language processing, machine learning, AIGC, and artificial intelligence.

Paper

arxiv.org/abs/2406.04325