datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Open-Sora-Plan-v1.1.0
Annotation
We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match the video names. Please refer to this https://github.com/PKU-YuanGroup/Open-Sora-Plan/issues/312#issuecomment-2197312973
Pexels
Pexels consists of multiple folders, but each folder exceeds the size limit for Huggingface uploads. Therefore, we divided each folder into 5 parts. You need to merge the 5 parts of each folder first, and then extract each… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.1.0.UniWorld-V1
The Geneval-style dataset is sourced from BLIP3o-60k.
This dataset is presented in the paper: UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
More details can be found in UniWorld-V1
Data preparation
Download the data from LanguageBind/UniWorld-V1. The dataset consists of two parts: source images and annotation JSON files.
Prepare a data.txt file in the following format:
The first column is the root path to the image.
The second… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/UniWorld-V1.Open-Sora-Plan-v1.0.0
Open-Sora-Dataset
Welcome to the Open-Sora-DataSet project! As part of the Open-Sora-Plan project, we specifically talk about the collection and processing of data sets. To build a high-quality video dataset for the open-source world, we started this project. 💪
We warmly welcome you to join us! Let's contribute to the open-source world together! Thank you for your support and contribution.
If you like our project, please give us a star ⭐ on GitHub for latest update.… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.0.0.MoE-LLaVA
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
If you like our project, please give us a star ⭐ on GitHub for latest update.
📰 News
[2024.01.30] The paper is released.
[2024.01.27] 🤗Hugging Face demo and all codes & datasets are available now! Welcome to watch 👀 this repository for the latest updates.
😮 Highlights
MoE-LLaVA shows excellent performance in multi-modal learning.
🔥 High performance, but with fewer… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/MoE-LLaVA.Video-LLaVAVideo-BenchCambrian737kOpen-Sora-Plan-v1.2.0
10M SAM
The original json was obtained from v1.1.0, just with the RESOLUTION information added.
The format of image annotation file is as follows.
[
{
"path": "00168/001680102.jpg",
"cap": [
"xxxxx."
],
"resolution": {
"height": 512,
"width": 683
}
},
...
]
6M HQ Panda70m
The format of video annotation file is as follows. Each element's path follows the structure: part_x/youtube_id/youtube_id_segment_i.mp4.
Here, part_x is… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.2.0.VIDAL-Depth-Thermal
【ICLR 2024 🔥】LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
If you like our project, please give us a star ⭐ on GitHub for latest update.
📰 News
[2024.01.27] 👀👀👀 Our MoE-LLaVA is released! A sparse model with 3B parameters outperformed the dense model with 7B parameters.
[2024.01.16] 🔥🔥🔥 Our LanguageBind has been accepted at ICLR 2024! We earn the score of 6(3)8(6)6(6)6(6) here.
[2023.12.15] 💪💪💪 We… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/VIDAL-Depth-Thermal.Open-Sora-Plan-v1.3.0We have open-sourced our dataset of 32,555 pairs, which includes Chinese data. The dataset is available here. The details can be found here.
In fact, it is a JSON file with the following structure. More details can be found here.
[
{
"instruction": "Refine the sentence: \"A newly married couple sharing a piece of there wedding cake.\" to contain subject description, action, scene description. (Optional: camera language, light and shadow, atmosphere) and conceive some additional actions… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.3.0.LLMBindLLMBind-GPT-Interactive-DataDreamDataStyleVideoDataSetdream2.tar.gz
