CoolFace
Modelpublic

open-gigaai/GigaWorld-0-Video-GR1-2b

sourceHugging Faceapache-2.0updated 10mo agoView on Hugging Face
7likes
Model Card

<div align="center" style="font-family: charter;"> <h1> GigaWorld-0: World Models as Data Engine to Empower Embodied AI </h1>

![License](https://opensource.org/licenses/Apache-2.0) ![Project](https://gigaworld0.github.io/) ![Papers](https://arxiv.org/abs/2511.19861) ![Demo](https://github.com/gigaworld0/gigaworld0.github.io/releases/download/v1/Gigaworld-0.mp4) ![Code](https://github.com/open-gigaai/giga-world-0/tree/main)

</div>

✨ Introduction

World models are emerging as a foundational paradigm for scalable, data-efficient embodied AI. In this work, we present GigaWorld-0, a unified world model framework designed explicitly as a data engine for Vision-Language-Action (VLA) learning. GigaWorld-0 integrates two synergistic components: GigaWorld-0-Video, which leverages large-scale video generation to produce diverse, texture-rich, and temporally coherent embodied sequences under fine-grained control of appearance, camera viewpoint, and action semantics; and GigaWorld-0-3D, which combines 3D generative modeling, 3D Gaussian Splatting reconstruction, physically differentiable system identification, and executable motion planning to ensure geometric consistency and physical realism.

πŸ—ΊοΈ Architecture

GigaWorld-0-Video-Dreamer is our foundation video generation model, capable of achieving IT2V generation.

Dreamer

πŸ“– Citation

If you use GigaWorld-0 in your research, please cite:

bibtex
@misc{gigaai2025gigaworld0,
      title={GigaWorld-0: World Models as Data Engine to Empower Embodied AI},
      author={GigaAI},
      year={2025},
      eprint={2511.19861},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2511.19861},
}