datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Open-Sora-Plan-v1.1.0
Annotation
We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match the video names. Please refer to this https://github.com/PKU-YuanGroup/Open-Sora-Plan/issues/312#issuecomment-2197312973
Pexels
Pexels consists of multiple folders, but each folder exceeds the size limit for Huggingface uploads. Therefore, we divided each folder into 5 parts. You need to merge the 5 parts of each folder first, and then extract each… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.1.0.OpenVid-1M-wds
OpenVid-1M — WebDataset repackaging
This repository is a sequential-read-optimized WebDataset repackaging of nkp37/OpenVid-1M by Nan et al. (ICLR 2025). The video content is identical to the original — only the on-disk layout is changed so it can be streamed efficiently from a single HTTP/NFS connection.
What differs from the original
Aspect
Original nkp37/OpenVid-1M
This repository
Format
Per-video mp4 files zipped
WebDataset .tar shards (~2 GB each)… See the full description on the dataset page: https://huggingface.co/datasets/Dev-Jahn/OpenVid-1M-wds.open-pmc-18m
OPEN-PMC
Arxiv: Arxiv
|
Code: Open-PMC Github
|
Model Checkpoint: Hugging Face
Dataset Summary
This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes:
Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc-18m.Open-Sora-Plan-v1.0.0
Open-Sora-Dataset
Welcome to the Open-Sora-DataSet project! As part of the Open-Sora-Plan project, we specifically talk about the collection and processing of data sets. To build a high-quality video dataset for the open-source world, we started this project. 💪
We warmly welcome you to join us! Let's contribute to the open-source world together! Thank you for your support and contribution.
If you like our project, please give us a star ⭐ on GitHub for latest update.… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.0.0.OpenSubject
OpenSubject Dataset
OpenSubject is a video-derived large-scale corpus with 2.5M samples and 4.35M images for subject-driven generation and manipulation, as presented in the paper OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation.
Project Page & Code
See the main repository for more details and code: OpenSubject
Dataset Structure
OpenSubject/
├── Images_packages/ # Compressed image… See the full description on the dataset page: https://huggingface.co/datasets/AIPeanutman/OpenSubject.openve_subSimScale
Haochen Tian,
Tianyu Li,
Haochen Liu,
Jiazhi Yang,
Yihang Qiu,
Guang Li,
Junli Wang,
Yinfeng Gao,
Zhang Zhang,
Liang Wang,
Hangjun Ye,
Tieniu Tan,
Long Chen,
Hongyang Li
📧 Primary Contact: Haochen Tian (tianhaochen2023@ia.ac.cn)
📜 Materials: 🌐 𝕏 | 📰 Media| 🗂️ Slides | 🎬 Talk (in Chinese)
🖊️ Joint effort by CASIA, OpenDriveLab at HKU, and Xiaomi EV.
🔥 Highlights
🏗️ A scalable simulation pipeline that synthesizes diverse and… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab-org/SimScale.open_video_dataOpenDialog
OpenDialog
OpenDialog is a 6.8k hours spoken dialogue dataset, introduced in the paper ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching.
Paper: https://arxiv.org/abs/2507.09318
GitHub: https://github.com/k2-fsa/ZipVoice
Project Page: https://zipvoice-dialog.github.io
OpenDialog is the first large-scale (6.8k hours) open-source spoken dialogue dataset derived from in-the-wild speech data. It consists of:
English data: 5074 hours
Chinese data: 1759… See the full description on the dataset page: https://huggingface.co/datasets/k2-fsa/OpenDialog.open-pmc
OPEN-PMC
Arxiv: Arxiv
|
Code: Open-PMC Github
|
Model Checkpoint: Hugging Face
Dataset Summary
This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes:
Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc.Open-PPIOpen-Sora-Plan-v1.0.0Open Sora plan collected 40,258 high-quality, watermark-free videos from open-source websites under the CC0 license. About 60% of the videos are in landscape format, with a total duration of approximately 274 hours, 5 minutes, and 13 seconds.
The dataset is divided into three main sources:
Mixkit:
Videos: 1,234
Total duration: 6h 19m 32s
Total frames: 570,815
Resolution and aspect ratio distributions (less than 1% not listed).
Pexels:
Videos: 7,408
Total duration: 48h 49m 24s
Total… See the full description on the dataset page: https://huggingface.co/datasets/Hemgg/Open-Sora-Plan-v1.0.0.OpenS2S_Datasets
How to Use?
Download, merge the files, and extract
You can run the following command to merge the compressed file parts after downloading.
cat en_response_wav.tar.gz.* > en_response_wav.tar.gz
cat zh_response_wav.tar.gz.* > zh_response_wav.tar.gz
OpenDriveLab___OpenLane
简介
OpenLane 是迄今为止第一个真实世界和规模最大的 3D 车道数据集。我们的数据集从公共感知数据集 Waymo Open Dataset 中收集有价值的内容,并为 1000 个路段提供车道和最近路径对象(CIPO)注释。简而言之,OpenLane 拥有 200K 帧和超过 880K 仔细注释的车道。我们公开发布了 OpenLane 数据集,以帮助研究界在 3D 感知和自动驾驶技术方面取得进步。
类定义
Lane Annotation
Lane shape. Each 2D/3D lane is presented as a set of 2D/3D points.
Lane category. Each lane has a category such as double yellow line or curb.
Lane property. Some of lanes have a property such as right, left.
Lane tracking ID. Each lane except… See the full description on the dataset page: https://huggingface.co/datasets/AlayaNeW/OpenDriveLab___OpenLane.OpenVid-60k-split
Combination of part_id's from bigdata-pw/OpenVid-1M and video data from nkp37/OpenVid-1M.
This is a 60k video split of the original dataset for faster iteration during testing. The split was obtained by filtering on aesthetic and motion scores by iteratively increasing their values until there were at most 1000 videos. Only videos containing between 80 and 240 frames were considered.
from datasets import load_dataset, disable_caching, DownloadMode
from torchcodec.decoders import… See the full description on the dataset page: https://huggingface.co/datasets/finetrainers/OpenVid-60k-split.OpenVid-10k-split
Combination of part_id's from bigdata-pw/OpenVid-1M and video data from nkp37/OpenVid-1M.
This is a 10k video split of the original dataset for faster iteration during testing. The split was obtained by filtering on aesthetic and motion scores by iteratively increasing their values until there were at most 1000 videos. Only videos containing between 80 and 240 frames were considered.
from datasets import load_dataset, disable_caching, DownloadMode
from torchcodec.decoders import… See the full description on the dataset page: https://huggingface.co/datasets/finetrainers/OpenVid-10k-split.rechtspraak-opendata
Rechtspraak OpenData Archive
This dataset repository stores archival snapshots used by the Maastricht Rechtspraak data pipelines. The repository is organized as source-oriented data, not as a normalized analysis dataset.
Repository Layout
raw_data/ contains raw Rechtspraak OpenData source snapshots as downloaded from the upstream OpenData feeds.
exports/ contains compressed export snapshots produced by the rs-migration pipeline.
Current Contents… See the full description on the dataset page: https://huggingface.co/datasets/davidwickerhf/rechtspraak-opendata.Swift-OpenX-Embodiment-wrist-imagesFRoM-W1-Datasets
FRoM-W1: Towards General Humanoid Whole-Body Control with Language Instructions
The Humanoid Intelligence Team from FudanNLP and OpenMOSS
Introduction
Humanoid robots are capable of performing various actions such as greeting, dancing and even backflipping. However, these motions are often hard-coded or specifically trained, which limits their versatility. In this work, we present FRoM-W1[^1]… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/FRoM-W1-Datasets.open-lm-instruction-dataopen-sora-pexels-subset
Open-Sora Pexels Dataset (Captioned Only)
A curated subset of the Pexels videos from LanguageBind/Open-Sora-Plan-v1.1.0, converted to layered WebDataset format. Every video has at least one caption.
Dataset Summary
Statistic
Value
Total Videos
9,750
Total Caption Entries
31,910
Captions from 513f source
4,452
Captions from 65f source
27,458
Video Shards
~120
Total Size
~120 GB
Caption Sources
Captions are merged from two Open-Sora… See the full description on the dataset page: https://huggingface.co/datasets/zengxianyu/open-sora-pexels-subset.OpenWebText2
Dataset Card for OpenWebText2
OpenWebText2 is a reasonably large corpus of scraped natural language data.
Original hosting for this dataset has become difficult because it was hosted alongside another controversial dataset. To the best of my knowledge, this dataset itself is not encumbered in any way. It's a useful size for smaller language modelling experiments and is sometimes used in existing papers which it may be desirable to replicate. It is uploaded here to facilitate those… See the full description on the dataset page: https://huggingface.co/datasets/segyges/OpenWebText2.OpenVid384pxOpen-Qwen2VL-Data-InterleavedYouTube-Commons-5G-RawOpenVid-1k-split
Combination of part_id's from bigdata-pw/OpenVid-1M and video data from nkp37/OpenVid-1M.
This is a 1k video split of the original dataset for faster iteration during testing. The split was obtained by filtering on aesthetic and motion scores by iteratively increasing their values until there were at most 1000 videos. Only videos containing between 80 and 240 frames were considered.
Loading the data:
from datasets import load_dataset, disable_caching, DownloadMode
from… See the full description on the dataset page: https://huggingface.co/datasets/finetrainers/OpenVid-1k-split.Youtube-Common-First-600openasl-dwpose
Dataset
This is a derivative of the OpenASL dataset, but contains pose estimations instead of the original RGB videos.Pose keypoints are predicted using DWPose, and each sample is stored as a pickle file.All the keypoints have been generated using the original videos in their full resolutions to maximize accuracy.
Field
Type
Description
poses
list
Pose data inferred from the video frames by DWPose
video_frames_per_second
float
Frame rate of the video
video_duration… See the full description on the dataset page: https://huggingface.co/datasets/PladsElsker/openasl-dwpose.mid-classical-openmusenet4-3mopen-pmc-18m
OPEN-PMC
Arxiv: Arxiv
|
Code: Open-PMC Github
|
Model Checkpoint: Hugging Face
Dataset Summary
This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes:
Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/chjudy/open-pmc-18m.
