CoolFace
Apppublic

lulavc/Z-Image-Turbo

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
31likes
App README

⚡ Z-Image Turbo

Ultra-fast text-to-image generation powered by [Tongyi-MAI/Z-Image-Turbo](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) — a single-stream Diffusion Transformer that produces stunning, high-quality images in just 8 steps.


✨ Features

FeatureDetails
🚀 8-step generationUltra-fast inference — typically completes in seconds on ZeroGPU
🎨 Single-stream DiTZ-Image architecture: efficient, high-quality single-stream Diffusion Transformer
⚡ AoTI accelerationAhead-of-Time Inductor compiled blocks (FA3 → standard fallback)
🎯 Quality PresetsOne-click presets: Fast, Balanced (default), Quality, Maximum — optimize speed vs detail
✨ AI Prompt EnhancerOne-click prompt enrichment via LLM — adds detail, lighting, composition
📐 13 resolution presetsStandard (1024-range) and Large (1280+) presets across all aspect ratios
🌐 4 languagesFull UI in English, Português (BR), Español, and عربي (Egyptian Arabic)
🎲 Smart seedingRandom or fixed seed — copy seed to reproduce any image exactly
🚫 Negative promptsFull negative prompt support to steer away from unwanted elements
🤖 MCP serverUse as a tool in Claude, Cursor, and any MCP-compatible AI agent
📥 Download & ShareOne-click download or share generated images directly from the UI

🖼️ Resolution Presets

Standard (1024-range)

PresetDimensionsAspect RatioBest for
◻ 1:11024 × 1024SquareSocial media, icons, avatars
▬ 4:31152 × 864LandscapeGeneral photography
▮ 3:4864 × 1152PortraitPeople, fashion
▬ 16:91280 × 720WidescreenWallpapers, cinematic
▮ 9:16720 × 1280VerticalMobile wallpapers, Reels
▬ 3:21248 × 832PhotoDSLR-style photography
▮ 2:3832 × 1248Photo portraitBook covers, posters
▬ 21:91344 × 576UltrawideCinematic banners

Large (1280+)

PresetDimensionsAspect Ratio
◻ 1:11280 × 1280Square
▬ 3:21536 × 1024Landscape
▮ 2:31024 × 1536Portrait
▬ 16:91536 × 864Widescreen
▮ 9:16864 × 1536Vertical
Note: Large presets require more VRAM. If you get an out-of-memory error, switch to a Standard preset or reduce the number of steps.

🌐 Language Support

The full UI (labels, placeholders, buttons, error messages) is available in 4 languages. Switch with the language selector at the top of the control panel:

FlagLanguageExample Prompts
🇺🇸EnglishAstronaut, jazz musician, epic fantasy, snowy cabin
🇧🇷Português (Brasil)Salvador market, Amazon rainforest, Rio samba, underwater city
🇪🇸EspañolGuatemala market, futuristic city, flamenco Sevilla, Mayan temple
🇪🇬عربي (مصري)Cairo woman, Alexandria aerial, Mamluk knight, Andalusian palace

You can write prompts in any language — the model handles multilingual input natively.


✨ AI Prompt Enhancer

Click the ✨ Enhance Prompt button to automatically improve your prompt using a large language model.

What it does:

  • —Locks in your core subject and intent
  • —Adds professional lighting, composition, and material details
  • —Applies cinematic or photographic aesthetics
  • —Handles text elements with precise quotation formatting

Example:

  • —Input: "a woman in a red dress"
  • —Enhanced: "Cinematic close-up portrait of a woman in a flowing crimson silk dress. Soft rim lighting from the left, deep shadow on the right side of her face. Shallow depth of field, f/1.8, bokeh background of blurred warm golden lights. High fashion editorial style, Leica M10 film grain."

The enhancer uses Qwen/Qwen2.5-7B-Instruct via the HF Inference API. Generation continues normally even if enhancement is unavailable.


🎯 Quality Presets

Quick presets to balance speed and quality:

PresetGuidance ScaleStepsBest For
⚡ Fast0.08Maximum speed, good for quick iterations
⚖️ Balanced (default)1.08Optimal balance of speed and quality
🎨 Quality3.58Better detail and prompt adherence, natural colors
💎 Maximum5.010Highest quality without oversaturation
Note: Z-Image Turbo is optimized for lower guidance scales than base models. Values above 5.0 may cause oversaturation and blown highlights.

⚙️ Advanced Settings

SettingDefaultRangeDescription
Steps81–30Number of diffusion steps. 8 = 7 DiT forward passes. More steps = slightly more detail but slower.
Guidance Scale1.00–10CFG scale. 0 = fastest; 1 = balanced (recommended); 3-5 = better quality; 7-10 = maximum detail. Higher values make the model follow your prompt more strictly.
Time Shift3.01–10Flow matching scheduler shift parameter. Controls frequency emphasis during sampling. Lower = more low-frequency structure, higher = more high-frequency detail.
Negative Prompt(empty)—Describe what you do NOT want in the image. E.g. blur, watermark, deformed hands.
SeedRandom0–2³²Set a fixed seed to reproduce the same image. The used seed is displayed after generation.

🤖 MCP Server

This Space runs as an MCP (Model Context Protocol) server, allowing it to be used as a tool by AI agents.

Using with Claude Desktop

Add to your claude_desktop_config.json:

json
{
  "mcpServers": {
    "z-image-turbo": {
      "url": "https://lulavc-z-image-turbo.hf.space/gradio_api/mcp/sse"
    }
  }
}

Using with any MCP client

  • —SSE endpoint: https://lulavc-z-image-turbo.hf.space/gradio_api/mcp/sse
  • —Tool name: generate
  • —Parameters: prompt, negative_prompt, resolution, steps, guidance_scale, shift, seed, random_seed

🔧 Technical Details

Model Architecture

  • —Name: Tongyi-MAI/Z-Image-Turbo
  • —Architecture: Z-Image — single-stream Diffusion Transformer (DiT)
  • —Scheduler: FlowMatch Euler Discrete with configurable time shift
  • —Precision: bfloat16
  • —License: Apache 2.0

Acceleration Stack

AoTI FA3 (flash-attn3 kernel)
    ↓ fallback
AoTI standard (compiled transformer blocks)
    ↓ fallback
Standard PyTorch inference

spaces.aoti_blocks_load pre-compiles the transformer blocks using Torch's Ahead-of-Time Inductor, significantly reducing per-step latency. The space automatically uses the best available kernel.

Infrastructure

  • —Hardware: ZeroGPU — NVIDIA H200 (141 GB HBM3e)
  • —Framework: Gradio 6.0.2
  • —Scheduler cache: Shift-keyed cache avoids re-instantiation on every call
  • —GPU memory: torch.cuda.empty_cache() called after every generation

💡 Tips for Best Results

  1. 1.Be specific — describe lighting, style, mood, camera angle, and materials
  2. 2.Name a style — "cinematic", "documentary", "concept art", "Studio Ghibli", "Leica film grain"
  3. 3.Use the Enhancer — click ✨ Enhance Prompt before generating for dramatically richer output
  4. 4.Start with Balanced preset — default guidance scale 1.0 provides optimal quality/speed balance
  5. 5.Use Quality preset for important work — guidance scale 3.5 significantly improves detail without oversaturation
  6. 6.Maximum preset for final output — guidance scale 5.0 with 10 steps provides highest quality
  7. 7.Avoid high guidance scales — values above 5.0 can cause oversaturation and color artifacts in Turbo models
  8. 8.Time Shift 3 — good default; try higher (5–7) for more detailed textures
  9. 9.Fix your seed — uncheck Random Seed, copy the seed after a good generation, reuse it with tweaks
  10. 10.Negative prompts — blur, deformed hands, watermark, low quality work well as defaults

🛠️ Running Locally

bash
git clone https://huggingface.co/spaces/lulavc/Z-Image-Turbo
cd Z-Image-Turbo
pip install -r requirements.txt
python app.py

Requires a CUDA-capable GPU with at least 16 GB VRAM for 1024×1024. Set HF_TOKEN in your environment for prompt enhancement.


📄 License


Space by [lulavc](https://huggingface.co/lulavc) · Powered by Tongyi-MAI/Z-Image-Turbo · ZeroGPU · A10G