CoolFace
Modelpublic

Zhoues/MineDreamer-InstructPix2Pix-Unet

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
2likes
Model Card

Model Card for MineDreamer 🔥

<!-- Provide a quick summary of what the model is/does. -->

![arXiv](https://arxiv.org/abs/2403.12037)

![project page](https://sites.google.com/view/minedreamer/main)

MineDreamer is an instructable embodied agent for simulated control and it is developed on top of recent advances in Multimodal Large Language Models (MLLMs) and diffusion models!

<p align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/63f08dc79cf89c9ed1bb89cd/S62I1Tn5qz5qJ3IkgMHH8.png" width=93%> <p>

MineDreamer can follow instructions steadily by employing a Chain-of-Imagination (CoI) mechanism to envision the step-by-step process of executing instructions and translating imaginations into more precise visual prompts tailored to the current state; subsequently, it generates keyboard-and-mouse actions to efficiently achieve these imaginations,

<p align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/63f08dc79cf89c9ed1bb89cd/LJxBMChCFng_RkXwUotfk.png" width=93%> <p>

This repo is used for hosting MineDreamer's InstructPix2Pix checkpoints, which are not only the baseline checkpoints but the training stage 2 checkpoints for Imaginator as well.

For more details or tutorials see https://github.com/Zhoues/MineDreamer.