CoolFace
Modelpublic

cocxxe/bonsai-image-ternary-4B-gemlite-2bit

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes1downloads
Model Card

<p align="center"> <img src="./assets/bonsai-logo.svg" width="280" alt="Bonsai Image"> </p>

<p align="center"> <a href="https://prismml.com"><b>Prism ML Website</b></a> &nbsp;|&nbsp; <a href="https://github.com/PrismML-Eng/Bonsai-Image-Demo/blob/main/bonsai-image-4b-whitepaper.pdf"><b>White Paper</b></a> &nbsp;|&nbsp; <a href="https://github.com/PrismML-Eng/Bonsai-Image-Demo"><b>Demo &amp; Examples</b></a> &nbsp;|&nbsp; <a href="https://discord.gg/prismml"><b>Discord</b></a> </p>

bonsai-image-ternary-4B-gemlite-2bit

Ternary weight (1.58-bit) text-to-image diffusion transformer deployment for NVIDIA GPUs

1.21 GB transformer | 6.4× smaller than FP16 | 4.5 s / 1024² on RTX 3080 | 2.8 s / 1024² on A100 | runs natively on Linux and Windows

Highlights

  • —1.21 GB diffusion transformer, down from 7.75 GB for the FP16 FLUX.2 Klein 4B transformer
  • —Ternary {-1, 0, +1} transformer weights with FP16 group-wise scaling in the matrix-heavy transformer layers (Q/K/V projections, output projections, MLP weights)
  • —Quality-oriented Bonsai Image variant: the additional zero state improves visual quality and prompt fidelity while keeping the transformer compact
  • —4.55 GB CUDA deployment payload including the 4-bit text encoder and FP16 VAE — text encoder is offloaded after prompt encode, so the denoising loop only keeps the compact transformer and VAE resident
  • —4-step FlowMatch-Euler sampler with guidance = 1.0 and shift = 3.0 — no CFG, no negative prompts needed
  • —Gemlite low-bit GEMM path for NVIDIA GPUs, with HQQ used for the compressed text encoder
  • —Runs on Linux and Windows natively through the same CUDA / Gemlite deployment stack
  • —Cross-platform companion: also available as MLX 2-bit for Apple Silicon

Resources

  • —[White Paper](https://github.com/PrismML-Eng/Bonsai-Image-Demo/blob/main/bonsai-image-4b-whitepaper.pdf) — full benchmarks, kernels, and memory analysis
  • —[Demo repo](https://github.com/PrismML-Eng/Bonsai-Image-Demo) — one-command setup for Mac / Linux / Windows
  • —[Discord](https://discord.gg/prismml) — community + support
  • —Kernels: gemlite (fused low-bit GEMM) · HQQ (low-bit quantization runtime) · triton-windows (Windows path)

Model Overview

ItemSpecification
Base architectureFLUX.2 Klein 4B (MMDiT diffusion transformer)
Parameters~4.0B (transformer trunk)
Blocks25 MMDiT blocks: 5 double-stream + 20 single-stream
SamplerFlowMatchEuler, 4 steps, guidance = 1.0, shift = 3.0
Text encoderQwen3-4B at 4-bit HQQ (≈ 2.84 GB CUDA payload, offloaded after prompt encode)
VAEFlux2 32-channel latent, tiled decode (128 px tiles)
Native resolution1024×1024 (also supports 512×512 and arbitrary multiples of 32)
Weight formatGemlite INT2 pack, ternary values + FP16 group-wise scales
Transformer size1.21 GB model-level Bonsai representation; 1.54 GB CUDA packed deployment size
Total payload4.55 GB CUDA deployment payload (transformer + 4-bit text encoder + FP16 VAE)
Ternary coverageAll 100 matmul-heavy linears in the 25 MMDiT blocks
PlatformsLinux x86_64 + Windows native on NVIDIA GPUs
LicenseApache 2.0

Ternary Weight Representation: 1.58-bit g128

Each ternary weight takes a value from {−1, 0, +1} with one shared FP16 scale per group of 128 weights:

text
w_i = scale_g * t_i,    t_i in {−1, 0, +1}

Ternary values carry log₂(3) ≈ 1.585 bits of information per weight. With one FP16 scale per group of 128, the effective storage is:

text
b_eff ≈ log2(3) + 16/128 ≈ 1.585 + 0.125 ≈ 1.71 bits/weight

This gives an idealized 9.4× reduction relative to FP16 for the ternary transformer layers. A small set of precision-sensitive supporting tensors remains in FP16, so the final Ternary Bonsai Image 4B diffusion transformer is 1.21 GB, a 6.4x reduction from the 7.75 GB FP16 FLUX.2 Klein 4B transformer.

The ternary representation is applied to the matrix-heavy transformer layers, including Q / K / V projections, output projections, MLP linears, and the double-stream add-K / Q / V linears. Supporting tensors (less than 5% of the total parameters) such as modulation streams, embedders, output norm, and output projection remain FP16 for image quality and stability.

The CUDA deployment uses a Gemlite INT2 packed format. Ternary values are stored in 2-bit slots, with the fourth code unused. The model-level Bonsai representation is 1.21 GB; the deployed CUDA pack is 1.54 GB on disk due to runtime packing and alignment overhead in the current Gemlite path.

Memory

FormatTransformer sizeReductionRatio
FP16 FLUX.2 Klein 4B7.75 GB—1.0×
Ternary Bonsai Image 4B1.21 GB84.4%6.4×

CUDA deployment:

ComponentSize
Gemlite INT2 diffusion transformer1.54 GB
HQQ 4-bit text encoder2.84 GB
FP16 VAE0.17 GB
Total payload4.55 GB

At runtime, the text encoder is offloaded after prompt encoding. During denoising, the repeated image-generation loop is dominated by the compact ternary diffusion transformer and active image-generation components rather than the full payload.

Peak HBM at 1024² on RTX 3080 is ~6.8 GiB end-to-end (transformer + VAE + activation memory).

Best Practices

  • —Sampler: FlowMatchEuler-discrete with 4 steps, guidance = 1.0, shift = 3.0. The model is designed for 4 steps; running more steps does not improve quality significantly and can introduce artifacts.
  • —Resolution: native 1024² is the design target. 512² works for quick previews.
  • —Aspect ratios: multiples of 32 are supported, including 832x1248 and 1248x832.
  • —Prompting: natural-language prompts. Negative prompts are not required.
  • —Runtime memory: the text encoder is offloaded after prompt encoding, so the denoising loop is memory-light.

Quickstart

Bonsai Studio (Linux / Windows)

The simplest path is the Bonsai Image Demo repo, which sets up the full Bonsai Studio (FastAPI backend + Next.js frontend) and selects gemlite automatically on Linux / Windows:

bash
git clone https://github.com/PrismML-Eng/Bonsai-Image-Demo.git
cd Bonsai-Image-Demo
./setup.sh
./scripts/download_model.sh           # ternary is the default
./scripts/serve.sh

On Windows (PowerShell):

powershell
Set-ExecutionPolicy -Scope CurrentUser RemoteSigned   # one-time
.\setup.ps1
.\scripts\download_model.ps1
.\scripts\serve.ps1

Python API (backend_gpu)

For inference without the studio frontend:

python
from backend_gpu.server import build_pipeline

pipe = build_pipeline(model_id="prism-ml/bonsai-image-ternary-4B-gemlite-2bit")
image = pipe(
    prompt="A bonsai tree in a quiet ceramic studio, soft morning light",
    num_inference_steps=4,
    guidance_scale=1.0,
    height=1024,
    width=1024,
).images[0]
image.save("bonsai.png")

Throughput (CUDA / gemlite)

Warmed wall-clock per image, 4 denoising steps, guidance = 1.0, matched prompts and sampler settings.

Platform512² (s)1024² (s)Notes
A100 (Colab)1.12.8Ampere datacenter (40 GB)
RTX PRO 6000 Blackwell (Colab)1.02.1NVIDIA Blackwell, 96 GB VRAM
RTX 3080 10 GB1.44.5Ampere consumer; 6.8 GiB peak HBM at 1024²
RTX 3060 6 GB (laptop)3.317.5Ampere mobile; memory-bound at 1024²

The sub-2-bit pack keeps generation viable on commodity GPUs at 1024². The RTX 3080 10 GB reaches 4.5 s/image, while the 6 GB laptop RTX 3060 is the memory-constrained tail.

Benchmarks

Evaluated with matched generation settings across the comparison set on H100. GenEval uses the official 512x512 protocol. For HPSv3 and DPG-Bench, larger-backbone rows are evaluated at 1024x1024, while smaller-backbone rows are evaluated at their native 512x512 setting. Higher is better for all three benchmarks.

ModelTransformer (GB)GenEvalHPSv3DPG-Bench
Bonsai Image · Ternary 4B1.210.72312.220.851
Bonsai Image · Binary 4B0.930.67111.150.822
FLUX.2 Klein 4B7.750.81912.840.853
FLUX.1-schnell23.80.71612.670.848
SDXL5.140.30010.050.740
PixArt-Σ XL 21.200.54111.930.769
Stable Diffusion 1.51.720.3964.200.601
BK-SDM-Small0.980.2973.050.559

The benchmark results show the intended quality-footprint trade-off. Ternary Bonsai Image 4B is the quality-oriented variant: at 1.21 GB, it sits very close to FLUX.2 Klein 4B across GenEval, HPSv3, and DPG-Bench while reducing the diffusion transformer footprint by 6.4x. The binary companion is the footprint-oriented variant, reducing the diffusion transformer below 1 GB while still delivering strong benchmark results.

Together, the Bonsai Image variants move the quality-footprint frontier: they bring modern diffusion-transformer behavior into a memory range previously occupied by much smaller, lower-capability models.

Use Cases

  • —Local creative tooling: image generation directly on CUDA-equipped workstations and consumer GPUs
  • —Private generation: prompts and generated assets can remain in local or controlled environments
  • —Rapid iteration: lower local latency and no remote queue for iterative creative workflows
  • —Commodity-GPU serving: lower transformer footprint and reduced memory pressure for serving on NVIDIA GPUs
  • —Windows and Linux deployment: native paths through the same Gemlite deployment stack
  • —Enterprise and controlled inference: local or private environments for data residency and compliance-sensitive workflows

Limitations

  • —Ternary Bonsai Image 4B is not bit-identical to the FP16 FLUX.2 Klein 4B model; it is a compact ternary-weight deployment designed to deliver similar practical behavior at much smaller size.
  • —Image-generation quality remains prompt- and workflow-dependent. Small text, fine details, object counts, and strict compositional constraints should be evaluated for the target use case.
  • —Current commodity inference stacks do not yet expose fully native ternary execution as a standard hardware path. This release uses practical Gemlite low-bit GEMM kernels on CUDA.
  • —After the diffusion transformer is made compact, other components such as the VAE can become more visible memory bottlenecks. The runtime mitigates this with text-encoder offload and tiled VAE decoding.

Citation

bibtex
@techreport{bonsaiimage4b,
    title   = {Bonsai Image 4B: Low-Bit Diffusion on Apple Silicon and Consumer GPUs},
    author  = {Prism ML},
    year    = {2026},
    month   = {May},
    url     = {https://prismml.com}
}

Contact

For questions, feedback, or collaboration inquiries: contact@prismml.com