CoolFace
Apppublic

aaditech/demo-chatgpt

sourceHugging Faceupdated 4y agoView on Hugging Face
0likes
App README

Visual ChatGPT

Visual ChatGPT connects ChatGPT and a series of Visual Foundation Models to enable sending and receiving images during chatting.

See our paper: <font size=5>Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models</font>

Visual ChatGPT Colab Support

You can run the colab notebook with following models text to image,ImageCaptioning,BLIP VQA,image to canny ![Open 2k image generation in Colab](https://colab.research.google.com/drive/1vhF4f3091h1cHZUh5QK7qByBHUDKbSWA?usp=sharing)

Demo

<img src="./assets/demo_short.gif" width="750">

System Architecture

<p align="center"><img src="./assets/figure.jpg" alt="Logo"></p>

Quick Start

# create a new environment
conda create -n visgpt python=3.8

# activate the new environment
conda activate visgpt

#  prepare the basic environments
pip install -r requirement.txt

# download the visual foundation models
bash download.sh

# prepare your private openAI private key
export OPENAI_API_KEY={Your_Private_Openai_Key}

# create a folder to save images
mkdir ./image

# Start Visual ChatGPT !
python visual_chatgpt.py

GPU memory usage

Here we list the GPU memory usage of each visual foundation model, one can modify `self.tools` with fewer visual foundation models to save your GPU memory:

Fundation ModelMemory Usage (MB)
ImageEditing6667
ImageCaption1755
T2I6677
canny2image5540
line2image6679
hed2image6679
scribble2image6679
pose2image6681
BLIPVQA2709
seg2image5540
depth2image6677
normal2image3974
Pix2Pix2795

Acknowledgement

We appreciate the open source of the following projects:

  • HuggingFace [[Project]](https://github.com/huggingface/transformers)
  • ControlNet [[Paper]](https://arxiv.org/abs/2302.05543) [[Project]](https://github.com/lllyasviel/ControlNet)
  • Stable Diffusion [[Paper]](https://arxiv.org/abs/2112.10752) [[Project]](https://github.com/CompVis/stable-diffusion)