CoolFace
14 results

mobile-o

Amshaker /Mobile-O-Post-Train Mobile-O Post-Training Data Unified Multimodal Post-Training · ~105K Quadruplet Samples 📌 Overview This dataset is used for Stage 3: Unified Multimodal Post-Training of Mobile-O, a unified multimodal model for on-device understanding and generation. The goal of this stage is to jointly improve both image generation and visual understanding through a multi-task objective using quadruplet samples. 📊 Dataset Format Each sample is a quadruplet consisting of:… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-Post-Train.imagetext-to-image1K<n<10K13 likes6.7k downloads7mo agoHugging FaceAmshaker /Mobile-O-Pre-Train Mobile-O Pre-Training Data Cross-Modal Alignment · 9M Text-Image Pairs 📌 Overview This dataset is used for Stage 1: Cross-Modal Alignment pre-training of Mobile-O, a unified multimodal model for on-device understanding and generation. The goal of this stage is to align the DiT diffusion decoder and Mobile Conditioning Projector (MCP) with the frozen VLM backbone using large-scale text-image pairs. 📊 Dataset Composition Source Samples Description… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-Pre-Train.imagetext-to-image10M<n<100M12 likes2.6k downloads7mo agoHugging FaceAmshaker /Mobile-O-SFT Mobile-O SFT Data Supervised Fine-Tuning · ~105K Curated Prompt-Image Pairs 📌 Overview This dataset is used for Stage 2: Supervised Fine-Tuning (SFT) of Mobile-O, a unified multimodal model for on-device understanding and generation. The goal of this stage is to improve image generation quality by fine-tuning on high-quality curated prompt-image pairs. 📊 Dataset Composition Source Samples Description BLIP3o 60K High-quality prompt-image pairs… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-SFT.imagetext-to-image1K<n<10K5 likes343 downloads7mo agoHugging Face