CoolFace
Modelpublic

FINAL-Bench/POCKET-Image-Zimage

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
42likes42downloads
Model Card

๐Ÿ–ผ๏ธ POCKET-Image-Zimage โ€” 4-bit (NF4) Z-Image for on-device

A 4-bit (NF4) quantized build of [Z-Image](https://huggingface.co/Tongyi-MAI/Z-Image) (Apache-2.0), packaged by VIDRAFT for low-VRAM, on-device image generation โ€” part of the POCKET line.

  • โ€”๐Ÿ“ฆ ~6 GB on disk (transformer + text encoder in NF4, VAE in fp16)
  • โ€”โšก Runs from ~8.6 GB VRAM (โ‰ˆ4.5 GB with CPU offload) โ€” vs 23.3 GB for bf16
  • โ€”๐ŸŽฏ ~2.7โ€“5ร— smaller footprint, quality on par with the bf16 base

Usage

python
import torch
from diffusers import ZImagePipeline   # or ZImageImg2ImgPipeline / ZImageInpaintPipeline

pipe = ZImagePipeline.from_pretrained(
    "FINAL-Bench/POCKET-Image-Zimage", torch_dtype=torch.bfloat16
).to("cuda")
img = pipe("a serene mountain lake at sunrise, photorealistic", num_inference_steps=20).images[0]
img.save("out.png")

Requires `bitsandbytes` (CUDA). Measured reload + generate peak: ~10.9 GB VRAM. For Apple Silicon / CPU, an optimum-quanto int8 build (~13.4 GB) is the portable option.

๐ŸŽจ The full POCKET-Image system

This repo hosts the quantized base model only. The headline character-perfect Korean & multilingual text feature is delivered by the POCKET-Image pipeline, not by these weights alone. Try the full system here:

  • โ€”๐ŸŽจ Studio (generate here): https://huggingface.co/spaces/FINAL-Bench/POCKET-Image-Studio

Base model: Tongyi-MAI/Z-Image (Apache-2.0) ยท Quantization: bitsandbytes NF4 ยท By VIDRAFT.

<!-- POCKET-FAMILY -->


๐Ÿงฉ The POCKET Family โ€” On-device AI by VIDRAFT

Big models, small hardware. No GPU, no cloud.

Models

Demos & tools (Spaces)

๐Ÿ“š Full POCKET collection

<!-- /POCKET-FAMILY -->