datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mobile-O-Post-Train
Mobile-O Post-Training Data
Unified Multimodal Post-Training · ~105K Quadruplet Samples
📌 Overview
This dataset is used for Stage 3: Unified Multimodal Post-Training of Mobile-O, a unified multimodal model for on-device understanding and generation.
The goal of this stage is to jointly improve both image generation and visual understanding through a multi-task objective using quadruplet samples.
📊 Dataset Format
Each sample is a quadruplet consisting of:… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-Post-Train.Mobile-O-Pre-Train
Mobile-O Pre-Training Data
Cross-Modal Alignment · 9M Text-Image Pairs
📌 Overview
This dataset is used for Stage 1: Cross-Modal Alignment pre-training of Mobile-O, a unified multimodal model for on-device understanding and generation.
The goal of this stage is to align the DiT diffusion decoder and Mobile Conditioning Projector (MCP) with the frozen VLM backbone using large-scale text-image pairs.
📊 Dataset Composition
Source
Samples
Description… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-Pre-Train.Mobile-O-SFT
Mobile-O SFT Data
Supervised Fine-Tuning · ~105K Curated Prompt-Image Pairs
📌 Overview
This dataset is used for Stage 2: Supervised Fine-Tuning (SFT) of Mobile-O, a unified multimodal model for on-device understanding and generation.
The goal of this stage is to improve image generation quality by fine-tuning on high-quality curated prompt-image pairs.
📊 Dataset Composition
Source
Samples
Description
BLIP3o
60K
High-quality prompt-image pairs… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-SFT.
