TULLUS/Xiaomi-Robotics-U0-4B-Sequence
<p align="center"><strong>TULLUS = I chose to work with this varient because it is ANY-TO-ANY ๐ค+๐ฆพ</strong></p><hr>
<div align="center"> <h1>Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Models</h1> <p> Xiaomi Robotics </p> <p> <a href="https://arxiv.org/abs/2607.11643"><img src="https://img.shields.io/badge/arXiv-2607.11643-b31b1b.svg" alt="Paper" /></a> <a href="https://robotics.xiaomi.com/xiaomi-robotics-u0.html"><img src="https://img.shields.io/badge/Project-Page-2e7d32.svg" alt="Project Page" /></a> <a href="https://huggingface.co/collections/XiaomiRobotics/xiaomi-robotics-u0"><img src="https://img.shields.io/badge/%F0%9F%A4%97-HuggingFace-ffd21e.svg" alt="Hugging Face" /></a> <a href="https://modelscope.cn/collections/XiaomiRobotics/Xiaomi-Robotics-U0"><img src="https://img.shields.io/badge/ModelScope-Collection-624aff.svg?logo=modelscope&logoColor=white" alt="ModelScope" /></a> <a href="https://github.com/XiaomiRobotics/Xiaomi-Robotics-U0/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-Apache2.0-blue.svg" alt="License" /></a> </p> </div>
<div align="center"> <img src="assets/architecture.png" alt="Xiaomi-Robotics-U0 model architecture with autoregressive generation and FlashAR acceleration." width="100%" /> </div>
<div align="center"> <img src="assets/illustrate.png" alt="Xiaomi-Robotics-U0 task examples across image generation, embodied scene generation, transfer, and video generation." width="100%" /> </div>
Xiaomi-Robotics-U0 exposes six public task types through one autoregressive framework:
News
- [September 2026] ๐ฅ Released Xiaomi-Robotics-U0-4B, Xiaomi-Robotics-U0-Sequence, and Xiaomi-Robotics-U0-4B-Sequence weights.
- [September 2026] ๐ป Open-sourced the FSDP training code.
- [July 2026] ๐ Released the Technical Report.
- [July 2026] ๐ฅ Released Xiaomi-Robotics-U0 and Xiaomi-Robotics-U0-FlashAR weights.
- [July 2026] ๐ป Inference code and scripts are now live!
Table of Contents
Model & Weights
Xiaomi-Robotics-U0, Xiaomi-Robotics-U0-4B, and Xiaomi-Robotics-U0-FlashAR support Scene Gen, Transfer, T2I, and X2I. The Sequence checkpoints support interleave_subtask and interleave_video with the eager backend.
Inference
The complete inference implementation, environment setup, configuration reference, command-line examples, and distributed inference instructions are available in `inference/README.md`.
The repository supports both eager execution and vLLM backends for AR and FlashAR inference. A Gradio demo is also provided for interactive T2I, X2I, Scene Gen, and Transfer workflows.
Training
Training code and environment configuration are reserved for `training/README.md` and will be added in a future release.
Citation
If you find this work useful, please cite:
@misc{li2026xiaomiroboticsu0,
title = {{Xiaomi-Robotics-U0}: Unified Embodied Synthesis with World Foundation Model},
author = {Xinghang Li and Jun Guo and Qiwei Li and Long Qian and Hang Lai and Yueze Wang and Hongyu Yan and Jiahang Cao and Xi Chen and Jingen Qu and Jiaxi Song and Nan Sun and Hanye Zhao and Futeng Liu and Wanli Peng and Heyun Wang and Yunhong Wang and Caoyu Xia and Jack Zhao and Diyun Xiang and Hangjun Ye and Heng Qu and Huaping Liu and Jason Li},
year = {2026},
eprint = {2607.11643},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.11643}
}