datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Open-Sora-Plan-v1.1.0
Annotation
We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match the video names. Please refer to this https://github.com/PKU-YuanGroup/Open-Sora-Plan/issues/312#issuecomment-2197312973
Pexels
Pexels consists of multiple folders, but each folder exceeds the size limit for Huggingface uploads. Therefore, we divided each folder into 5 parts. You need to merge the 5 parts of each folder first, and then extract each… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.1.0.OmniWorld[ICLR 2026] OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
🎉NEWS
[2026.3.21] 🔥 OmniWorld-Game with Metric Scale is now released! Check out our latest model Pi3X (an enhanced version of Pi3), which leverages this data to achieve better performance!
[2026.1.26] 🎉 OmniWorld was accepted by ICLR 2026!
[2026.1.7] Update OmniWorld-Game, release RH20T-Robot, RH20T-Human, Ego-Exo4D, EgoDex, Epic-Kitchens.
[2025.11.11] The OmniWorld is… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/OmniWorld.wds_objectnetOVO-Bench
OVO-Bench: How Far is Your Video-LLMs from Real-World Online VideO Understanding?
🔥🔥OVO-Bench is accepted by CVPR 2025!🔥🔥
Important Note: Current codebase is modified compared to our initial arXiv paper. We strongly recommend that any use of OVO-Bench should be based on current edition.
Introduction
🌟 Three distinct problem-solving modes
Backward Tracing: trace back to past events to answer the question.Real-Time Visual Perception:… See the full description on the dataset page: https://huggingface.co/datasets/JoeLeelyf/OVO-Bench.3d_optical_flow_droid
3D Optical Flow DROID Dataset
Processed DROID robotics dataset with optical flow and scene flow annotations.
Dataset Structure
Organized by lab, each trajectory in separate tar.gz archive:
IPRL/IPRL+2023-06-19+Mon_Jun_19_23:27:48_2023.tar.gz
CLVR/CLVR+2023-...tar.gz
... (15 labs, ~33K trajectories)
Each trajectory contains:
metadata.json - Trajectory metadata
trajectory.h5 - Robot state and actions
camera_left/, camera_right/ - Camera data
rgb/ - RGB images
depth/ -… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/3d_optical_flow_droid.obelics_seed2_tokensPart of the OBELISC data set, including 32 Million samples, please refer to dataset.py to use this data
OwnerMobile-O-Post-Train
Mobile-O Post-Training Data
Unified Multimodal Post-Training · ~105K Quadruplet Samples
📌 Overview
This dataset is used for Stage 3: Unified Multimodal Post-Training of Mobile-O, a unified multimodal model for on-device understanding and generation.
The goal of this stage is to jointly improve both image generation and visual understanding through a multi-task objective using quadruplet samples.
📊 Dataset Format
Each sample is a quadruplet consisting of:… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-Post-Train.OpenVid-1M-wds
OpenVid-1M — WebDataset repackaging
This repository is a sequential-read-optimized WebDataset repackaging of nkp37/OpenVid-1M by Nan et al. (ICLR 2025). The video content is identical to the original — only the on-disk layout is changed so it can be streamed efficiently from a single HTTP/NFS connection.
What differs from the original
Aspect
Original nkp37/OpenVid-1M
This repository
Format
Per-video mp4 files zipped
WebDataset .tar shards (~2 GB each)… See the full description on the dataset page: https://huggingface.co/datasets/Dev-Jahn/OpenVid-1M-wds.wds_objectnetreazonspeecholmoearth_pretrain_datasetThis is the pre-training dataset for training the OlmoEarth pre-trained remote sensing foundation models.
Documentation is on GitHub at https://github.com/allenai/olmoearth_pretrain/blob/main/docs/Pretraining-Dataset.md
The dataset is released under CC BY 4.0. It includes data from the following sources:
Sentinel-2 L2A imagery from the European Space Agency, available under the Copernicus Sentinel Data and Service Legal Notice
Sentinel-1 GRD IW vv+vh imagery from the European Space Agency… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmoearth_pretrain_dataset.X2Edit-Dataset
X2Edit
Introduction
X2Edit Dataset is a comprehensive image editing dataset that covers 14 diverse editing tasks and exhibits substantial advantages over existing open-source datasets including AnyEdit, HQ-Edit, UltraEdit, SEED-Data-Edit, ImgEdit and OmniEdit.
For the relevant data construction scripts, model training and inference scripts, please refer to X2Edit.
News
2025/09/16: We are about to release a dataset constructed by Qwen-Image and… See the full description on the dataset page: https://huggingface.co/datasets/OPPOer/X2Edit-Dataset.Mobile-O-Pre-Train
Mobile-O Pre-Training Data
Cross-Modal Alignment · 9M Text-Image Pairs
📌 Overview
This dataset is used for Stage 1: Cross-Modal Alignment pre-training of Mobile-O, a unified multimodal model for on-device understanding and generation.
The goal of this stage is to align the DiT diffusion decoder and Mobile Conditioning Projector (MCP) with the frozen VLM backbone using large-scale text-image pairs.
📊 Dataset Composition
Source
Samples
Description… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-Pre-Train.cc12m-recaptionedNuplan-OccupancyOla-DataThis repository contains the data presented in Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment.
Code: https://github.com/Ola-Omni/Ola
open-pmc-18m
OPEN-PMC
Arxiv: Arxiv
|
Code: Open-PMC Github
|
Model Checkpoint: Hugging Face
Dataset Summary
This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes:
Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc-18m.IRRISIGHT
IRRISIGHT
IRRISIGHT is a large-scale multimodal dataset to address water availability problems in agriculture. It is designed to support supervised and semi-supervised learning tasks related to agricultural water use monitoring.
Due to the space constraints, we uploaded the files across multiple repositories as follows:
To download Pennsylvania and Maryland, use the current repository (OBH30/IRRISIGHT).
To download Arizona, Arkansas, Florida, Georgia, New Jersey, North Carolina… See the full description on the dataset page: https://huggingface.co/datasets/OBH30/IRRISIGHT.sdoml-lite
SDOML-lite
SDOML-lite is a lightweight alternative to the SDOML dataset specifically designed for machine learning applications in solar physics, providing continuous full-disk images of the Sun with magnetic field and extreme ultraviolet data in several wavelengths. The data source is the Solar Dynamics Observatory (SDO) space telescope, a NASA mission that has been in operation since 2010.
NASA’s SDO mission has generated over 20 petabytes of high-resolution solar imagery… See the full description on the dataset page: https://huggingface.co/datasets/oxai4science/sdoml-lite.ophnet_3dOpen-Sora-Plan-v1.0.0
Open-Sora-Dataset
Welcome to the Open-Sora-DataSet project! As part of the Open-Sora-Plan project, we specifically talk about the collection and processing of data sets. To build a high-quality video dataset for the open-source world, we started this project. 💪
We warmly welcome you to join us! Let's contribute to the open-source world together! Thank you for your support and contribution.
If you like our project, please give us a star ⭐ on GitHub for latest update.… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.0.0.osu-beatmaps
osu! Beatmaps Dataset (WebDataset)
A collection of ranked/loved osu! beatmaps with audio and chart data, in WebDataset format.
Dataset Variants
Variant
Audio Format
Description
original
MP3/OGG/WAV
Full quality original audio files
compressed
64kbps Mono Opus
Compressed audio for smaller download
from datasets import load_dataset
# Load original audio variant
ds = load_dataset("project-riz/osu-beatmaps", "original", streaming=True)
# Load compressed… See the full description on the dataset page: https://huggingface.co/datasets/project-riz/osu-beatmaps.objaverse_rendering_setlaions_got_talent_embs_only
laions_got_talent Whisper Embeddings (Embeddings + Metadata Only)
This dataset contains Whisper embeddings (NPY) and metadata (JSON). The original audio files (MP3) are NOT included.
Embeddings computed with: mkrausio/EmoWhisper-AnS-Small-v0.1
Includes original audio: No
Includes metadata: Yes (JSON)
Includes embeddings: Yes (NPY)
Creation date: 2025-05-11
openve_subcorrectvla_dataOpenSubject
OpenSubject Dataset
OpenSubject is a video-derived large-scale corpus with 2.5M samples and 4.35M images for subject-driven generation and manipulation, as presented in the paper OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation.
Project Page & Code
See the main repository for more details and code: OpenSubject
Dataset Structure
OpenSubject/
├── Images_packages/ # Compressed image… See the full description on the dataset page: https://huggingface.co/datasets/AIPeanutman/OpenSubject.SimScale
Haochen Tian,
Tianyu Li,
Haochen Liu,
Jiazhi Yang,
Yihang Qiu,
Guang Li,
Junli Wang,
Yinfeng Gao,
Zhang Zhang,
Liang Wang,
Hangjun Ye,
Tieniu Tan,
Long Chen,
Hongyang Li
📧 Primary Contact: Haochen Tian (tianhaochen2023@ia.ac.cn)
📜 Materials: 🌐 𝕏 | 📰 Media| 🗂️ Slides | 🎬 Talk (in Chinese)
🖊️ Joint effort by CASIA, OpenDriveLab at HKU, and Xiaomi EV.
🔥 Highlights
🏗️ A scalable simulation pipeline that synthesizes diverse and… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab-org/SimScale.Light-Omni-Training
Light-Omni Training Dataset
This repository contains the training data used by Light-Omni, a multimodal
agent framework for reflexive video understanding with long-term memory.
Light-Omni uses memory-augmented multimodal streams to train adapters for
memory construction, response generation, and reaction/action control.
Links
Project page: https://clare-nie.github.io/Light-Omni/
Code: https://github.com/Clare-Nie/Light-Omni
Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ClareNie/Light-Omni-Training.
