datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Open-Sora-Plan-v1.1.0
Annotation
We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match the video names. Please refer to this https://github.com/PKU-YuanGroup/Open-Sora-Plan/issues/312#issuecomment-2197312973
Pexels
Pexels consists of multiple folders, but each folder exceeds the size limit for Huggingface uploads. Therefore, we divided each folder into 5 parts. You need to merge the 5 parts of each folder first, and then extract each… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.1.0.Open-Sora-Plan-v1.0.0
Open-Sora-Dataset
Welcome to the Open-Sora-DataSet project! As part of the Open-Sora-Plan project, we specifically talk about the collection and processing of data sets. To build a high-quality video dataset for the open-source world, we started this project. 💪
We warmly welcome you to join us! Let's contribute to the open-source world together! Thank you for your support and contribution.
If you like our project, please give us a star ⭐ on GitHub for latest update.… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.0.0.Open-Sora-Plan-v1.0.0Open Sora plan collected 40,258 high-quality, watermark-free videos from open-source websites under the CC0 license. About 60% of the videos are in landscape format, with a total duration of approximately 274 hours, 5 minutes, and 13 seconds.
The dataset is divided into three main sources:
Mixkit:
Videos: 1,234
Total duration: 6h 19m 32s
Total frames: 570,815
Resolution and aspect ratio distributions (less than 1% not listed).
Pexels:
Videos: 7,408
Total duration: 48h 49m 24s
Total… See the full description on the dataset page: https://huggingface.co/datasets/Hemgg/Open-Sora-Plan-v1.0.0.open-sora-pexels-subset
Open-Sora Pexels Dataset (Captioned Only)
A curated subset of the Pexels videos from LanguageBind/Open-Sora-Plan-v1.1.0, converted to layered WebDataset format. Every video has at least one caption.
Dataset Summary
Statistic
Value
Total Videos
9,750
Total Caption Entries
31,910
Captions from 513f source
4,452
Captions from 65f source
27,458
Video Shards
~120
Total Size
~120 GB
Caption Sources
Captions are merged from two Open-Sora… See the full description on the dataset page: https://huggingface.co/datasets/zengxianyu/open-sora-pexels-subset.loopwan-opensora-pilot-v1
LoopWan Open-Sora-Plan pilot
Status: completed bounded curation. Counts: {"long_audit": 22, "train": 2000, "val": 128}.
Fixed 320x480, timestamp sampling at 16 FPS; train/validation crops are real
contiguous 10-second shots, audit crops 20 seconds. Sources are disjoint and
captions are matched to pinned official annotations. See DATASET_REPORT.md for
filter thresholds, caption limitations and full provenance.
Official dataset revision: ab77293def393e6938f11a7bfd12163decfb9620.… See the full description on the dataset page: https://huggingface.co/datasets/Nicholas0228/loopwan-opensora-pilot-v1.videogen-rewardbench-opensora1-2Open-Sora-Plan-v1.2.0
10M SAM
The original json was obtained from v1.1.0, just with the RESOLUTION information added.
The format of image annotation file is as follows.
[
{
"path": "00168/001680102.jpg",
"cap": [
"xxxxx."
],
"resolution": {
"height": 512,
"width": 683
}
},
...
]
6M HQ Panda70m
The format of video annotation file is as follows. Each element's path follows the structure: part_x/youtube_id/youtube_id_segment_i.mp4.
Here, part_x is… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.2.0.tip-i2v-opensoraopen-sora-mini-hf-layered
Open-Sora Mini HF (Layered Format)
Converted from wangxingjun778/open-sora-mini-hf
Structure
├── videos/ # Video-only TAR shards (immutable)
│ ├── mixkit_shard_0000.tar
│ └── ...
├── annotations/ # Annotations (can add new versions)
│ └── captions_v1.parquet # Vision-generated captions
└── manifest.parquet # Index: sample_id -> shard mapping
Stats
Metric
Value
Total samples with captions
8,247… See the full description on the dataset page: https://huggingface.co/datasets/zengxianyu/open-sora-mini-hf-layered.Open-Sora-Plan-v1.3.0We have open-sourced our dataset of 32,555 pairs, which includes Chinese data. The dataset is available here. The details can be found here.
In fact, it is a JSON file with the following structure. More details can be found here.
[
{
"instruction": "Refine the sentence: \"A newly married couple sharing a piece of there wedding cake.\" to contain subject description, action, scene description. (Optional: camera language, light and shadow, atmosphere) and conceive some additional actions… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.3.0.v14-fake-opensora12open-sora-mini-hfOpen-Sora-Plan-v1.3.0v14-fake-opensoraopen-sora-pexels-fullSorawiz__Gemma-Creative-9B-Base-details
Dataset Card for Evaluation run of Sorawiz/Gemma-Creative-9B-Base
Dataset automatically created during the evaluation run of model Sorawiz/Gemma-Creative-9B-Base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sorawiz__Gemma-Creative-9B-Base-details.Sorawiz__Gemma-9B-Base-details
Dataset Card for Evaluation run of Sorawiz/Gemma-9B-Base
Dataset automatically created during the evaluation run of model Sorawiz/Gemma-9B-Base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sorawiz__Gemma-9B-Base-details.Open-Sora-PlanOpenSoraDataOmniCam-Sep-1-OpenSoraPlanOpenSora
