open-sora
Open-Sora-Plan-v1.1.0
Annotation
We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match the video names. Please refer to this https://github.com/PKU-YuanGroup/Open-Sora-Plan/issues/312#issuecomment-2197312973
Pexels
Pexels consists of multiple folders, but each folder exceeds the size limit for Huggingface uploads. Therefore, we divided each folder into 5 parts. You need to merge the 5 parts of each folder first, and then extract each… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.1.0.Open-Sora-Plan-v1.0.0
Open-Sora-Dataset
Welcome to the Open-Sora-DataSet project! As part of the Open-Sora-Plan project, we specifically talk about the collection and processing of data sets. To build a high-quality video dataset for the open-source world, we started this project. 💪
We warmly welcome you to join us! Let's contribute to the open-source world together! Thank you for your support and contribution.
If you like our project, please give us a star ⭐ on GitHub for latest update.… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.0.0.Open-Sora-Plan-v1.0.0Open Sora plan collected 40,258 high-quality, watermark-free videos from open-source websites under the CC0 license. About 60% of the videos are in landscape format, with a total duration of approximately 274 hours, 5 minutes, and 13 seconds.
The dataset is divided into three main sources:
Mixkit:
Videos: 1,234
Total duration: 6h 19m 32s
Total frames: 570,815
Resolution and aspect ratio distributions (less than 1% not listed).
Pexels:
Videos: 7,408
Total duration: 48h 49m 24s
Total… See the full description on the dataset page: https://huggingface.co/datasets/Hemgg/Open-Sora-Plan-v1.0.0.open-sora-pexels-subset
Open-Sora Pexels Dataset (Captioned Only)
A curated subset of the Pexels videos from LanguageBind/Open-Sora-Plan-v1.1.0, converted to layered WebDataset format. Every video has at least one caption.
Dataset Summary
Statistic
Value
Total Videos
9,750
Total Caption Entries
31,910
Captions from 513f source
4,452
Captions from 65f source
27,458
Video Shards
~120
Total Size
~120 GB
Caption Sources
Captions are merged from two Open-Sora… See the full description on the dataset page: https://huggingface.co/datasets/zengxianyu/open-sora-pexels-subset.loopwan-opensora-pilot-v1
LoopWan Open-Sora-Plan pilot
Status: completed bounded curation. Counts: {"long_audit": 22, "train": 2000, "val": 128}.
Fixed 320x480, timestamp sampling at 16 FPS; train/validation crops are real
contiguous 10-second shots, audit crops 20 seconds. Sources are disjoint and
captions are matched to pinned official annotations. See DATASET_REPORT.md for
filter thresholds, caption limitations and full provenance.
Official dataset revision: ab77293def393e6938f11a7bfd12163decfb9620.… See the full description on the dataset page: https://huggingface.co/datasets/Nicholas0228/loopwan-opensora-pilot-v1.videogen-rewardbench-opensora1-2
