CoolFace
Modelpublic

MeiGen-AI/GenEvolve

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
10likes90downloads
Model Card

<div align="center">

<img src="assets/logo_genevolve.png" alt="GenEvolve" width="160">

<h1>GenEvolve</h1>

<p><strong><em>Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation</em></strong></p>

<p> <a href="https://arxiv.org/abs/2605.21605"> <img alt="Paper" src="https://img.shields.io/badge/๐Ÿ“„Paper-arXiv:2605.21605-b31b1b"></a> <a href="https://ephemeral182.github.io/GenEvolve/"> <img alt="Project Page" src="https://img.shields.io/badge/๐ŸŒProject-Page-1f6feb"></a> <a href="https://github.com/MeiGen-AI/GenEvolve"> <img alt="Code" src="https://img.shields.io/badge/๐Ÿ’พGitHub-Code-181717"></a> <a href="https://huggingface.co/datasets/MeiGen-AI/GenEvolve-Data-Bench"> <img alt="Dataset" src="https://img.shields.io/badge/๐Ÿค—Dataset-GenEvolve--Data-FFD21E"></a> </p>

</div>

This repository hosts the GenEvolve agent policy โ€” a Qwen3-VL-8B-Instruct backbone fine-tuned and self-evolved into a tool-orchestrated image-generation agent. Given a user request, the agent issues web/image searches, retrieves visual references, activates internal generation knowledge, and emits an executable prompt-reference program z = (gen_prompt, reference_images) that drives any reference-conditioned downstream generator (Qwen-Image-Edit, Nano Banana Pro, ...).

<div align="center"> <img src="assets/teaser.jpg" alt="GenEvolve teaser" width="100%">

<p><em>The same trained agent policy paired with two reference-conditioned generators โŸถ<br> <strong>Qwen-Image-Edit (open)</strong> &nbsp;ยท&nbsp; <strong>Nano Banana Pro (strong)</strong></em></p> </div>


โœจ Highlights

  • โ€”Tool-orchestrated trajectories. The agent calls search, image_search, and query_knowledge (8 callable generation skills) before producing a final program z = (gen_prompt, reference_images).
  • โ€”Self-evolution with Visual Experience Distillation. Best-vs-worst trajectory pairs are distilled token-level into the deployed student. No runtime memory at inference.
  • โ€”Generator-transferable. The same trained policy works with both an open-source generator (Qwen-Image-Edit-2511) and a strong proprietary generator (Nano Banana Pro).

๐Ÿ“Š Headline Results

GenEvolve-Bench (KScore, held-out split)

MethodGeneratorKScoreKnowledge-Anch.Quality-Anch.
Qwen-Image (raw)Qwen-Image0.29870.23840.3768
Nano Banana Pro (raw)Nano Banana Pro0.52980.51600.5477
Gen-Searcher 8BQwen-Image-Edit-25110.34930.32930.3745
Gen-Searcher 8BNano Banana Pro0.54810.54720.5492
GenEvolve (Ours)Qwen-Image-Edit-25110.36630.34100.3990
GenEvolve (Ours)Nano Banana Pro0.57390.56690.5830

WISE Benchmark (WiScore, six knowledge categories)

ModelCulturalTimeSpaceBiologyPhysicsChemistry**Overall**
GPT-4o0.810.710.890.830.790.740.80
Gen-Searcher-8B + Qwen-Image0.800.710.820.760.740.750.77
Mind-Brush0.830.690.840.710.850.680.78
GenEvolve + Qwen-Image-Edit0.840.740.870.830.810.830.82

๐Ÿง  Method Overview

<p align="center"><img src="assets/overview.png" alt="GenEvolve method overview" width="92%"></p>

For a user request, the agent samples a multi-turn trajectory of tool calls before emitting the final prompt-reference program. The downstream generator then renders the image.


๐Ÿ–ผ๏ธ Visual Demos

<p align="center"><img src="assets/visual_comparison.png" alt="Qualitative comparison" width="100%"></p>

<p align="center"><sub>Qualitative comparison on representative cases. <span style="color:#D97706">Orange</span> marks external/uncommon knowledge requirements; <span style="color:#2563EB">blue</span> marks internal generation-knowledge requirements.</sub></p>

๐ŸŽจ Gallery โ€” paired with Nano Banana Pro

<p align="center"><img src="assets/gallery_nano.jpg" alt="GenEvolve + Nano Banana Pro gallery" width="100%"></p>

<p align="center"><sub>The same agent policy with Nano Banana Pro as the downstream renderer. Examples cover spatial layout, text rendering, quantity counting, attribute binding, anatomy/pose, creative transfer, material physics, and aesthetic drawing.</sub></p>

๐ŸŽจ Gallery โ€” paired with Qwen-Image-Edit (open)

<p align="center"><img src="assets/gallery_qwen.jpg" alt="GenEvolve + Qwen-Image-Edit gallery" width="100%"></p>

<p align="center"><sub>Same trained policy paired with the open-source Qwen-Image-Edit-2511 renderer; consistent quality across both generators reflects generator-transferable orchestration.</sub></p>


๐Ÿš€ Quick Start

The deployed checkpoint is the student policy โ€” it consumes a user prompt and returns a JSON gen_prompt + reference_images program through a <think>/<tool_call>/<answer> loop. The end-to-end runtime (vLLM serving + agent loop + tools + Qwen/Nano renderers) lives in the GitHub repo; the snippet below mirrors its installation and usage.

1. Install the main GenEvolve runtime

bash
git clone https://github.com/MeiGen-AI/GenEvolve.git
cd GenEvolve

conda create -n genevolve python=3.11 -y && conda activate genevolve
pip install -U pip setuptools wheel packaging psutil ninja
pip install torch==2.8.0 torchvision==0.23.0 --index-url https://download.pytorch.org/whl/cu128
pip install --no-build-isolation -r requirements.txt
pip install -e .

Qwen-Image-Edit rendering runs as a separate FastAPI service (kept out of the vLLM environment to avoid CUDA/diffusers conflicts). Set up that service from the GitHub README when you want to use --backend qwen-image-edit-service.

2. Serve the agent policy

bash
# Single GPU / single replica.
MODEL_PATH=MeiGen-AI/GenEvolve PORT=8000 TP=1 DP=1 bash scripts/serve_vllm.sh

# Higher throughput on one 8-GPU node (8 replicas, 1 GPU each).
MODEL_PATH=MeiGen-AI/GenEvolve PORT=8000 TP=1 DP=8 bash scripts/serve_vllm.sh

TP shards one model replica across multiple GPUs; DP launches multiple replicas; total GPU usage is TP ร— DP.

3. End-to-end example

bash
export SERPER_API_KEY=<your_key>      # required for search / image_search
export GOOGLE_API_KEY=<your_key>      # or GEMINI_API_KEY; only for --backend nano-banana-pro

# Nano Banana Pro renderer
python examples/quickstart.py \
    --backend nano-banana-pro \
    --base-url http://localhost:8000/v1 \
    --model GenEvolve \
    --prompt "A 1990s travel-magazine cover of two backpackers in front of the Eiffel Tower at golden hour, the title \"PARIS\" in bold serif." \
    --output paris.png

# Qwen-Image-Edit renderer (point at your Qwen-Image-Edit FastAPI service)
python examples/quickstart.py \
    --backend qwen-image-edit-service \
    --service-url http://your-qwen-service:8001 \
    --base-url http://localhost:8000/v1 \
    --model GenEvolve \
    --output paris_qwen.png

The agent's final <answer> is a JSON object:

json
{
  "gen_prompt": "...natural-language prompt that refers to images by 'the first reference image', ...",
  "reference_images": [
    {"img_id": "IMG_001", "note": "what to copy from this image"}
  ]
}

gen_prompt MUST refer to selected images using ordinal phrases ("the first reference image") โ€” never raw IMG_### ids or URLs. Pass (gen_prompt, [r["local_path"] for r in reference_images]) to your favourite reference-conditioned generator (Qwen-Image-Edit, Nano Banana Pro, ...) to obtain the final image.


๐Ÿ—‚๏ธ Related Artifacts

ArtifactLink
Project pagehttps://ephemeral182.github.io/GenEvolve/
PaperComing soon
Codehttps://github.com/MeiGen-AI/GenEvolve
Training data + benchmarkMeiGen-AI/GenEvolve-Data-Bench
Base modelQwen/Qwen3-VL-8B-Instruct

โš–๏ธ Intended Use, Limits, Bias

  • โ€”Intended use. Research on tool-using image-generation agents, agentic prompt-program synthesis, and self-distillation from generated outcomes.
  • โ€”Search dependency. The agent issues live web/image queries through user-provided tool wrappers. Quality of grounded facts depends on the search backend you plug in.
  • โ€”Bias. Tool outputs and reference images come from public web search, which carries demographic, cultural, and geographic biases that may be reflected in agent outputs.

๐Ÿ“‘ Citation

bibtex
@misc{chen2026genevolveselfevolvingimagegeneration,
      title={GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation}, 
      author={Sixiang Chen and Zhaohu Xing and Tian Ye and Xinyu Geng and Yunlong Lin and Jianyu Lai and Xuanhua He and Fuxiang Zhai and Jialin Gao and Lei Zhu},
      year={2026},
      eprint={2605.21605},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2605.21605}, 
}