cocxxe/bonsai-image-ternary-4B-gemlite-2bit
<p align="center"> <img src="./assets/bonsai-logo.svg" width="280" alt="Bonsai Image"> </p>
<p align="center"> <a href="https://prismml.com"><b>Prism ML Website</b></a> | <a href="https://github.com/PrismML-Eng/Bonsai-Image-Demo/blob/main/bonsai-image-4b-whitepaper.pdf"><b>White Paper</b></a> | <a href="https://github.com/PrismML-Eng/Bonsai-Image-Demo"><b>Demo & Examples</b></a> | <a href="https://discord.gg/prismml"><b>Discord</b></a> </p>
bonsai-image-ternary-4B-gemlite-2bit
Ternary weight (1.58-bit) text-to-image diffusion transformer deployment for NVIDIA GPUs
1.21 GB transformer | 6.4× smaller than FP16 | 4.5 s / 1024² on RTX 3080 | 2.8 s / 1024² on A100 | runs natively on Linux and Windows
Highlights
- 1.21 GB diffusion transformer, down from 7.75 GB for the FP16 FLUX.2 Klein 4B transformer
- Ternary {-1, 0, +1} transformer weights with FP16 group-wise scaling in the matrix-heavy transformer layers (Q/K/V projections, output projections, MLP weights)
- Quality-oriented Bonsai Image variant: the additional zero state improves visual quality and prompt fidelity while keeping the transformer compact
- 4.55 GB CUDA deployment payload including the 4-bit text encoder and FP16 VAE — text encoder is offloaded after prompt encode, so the denoising loop only keeps the compact transformer and VAE resident
- 4-step FlowMatch-Euler sampler with guidance = 1.0 and shift = 3.0 — no CFG, no negative prompts needed
- Gemlite low-bit GEMM path for NVIDIA GPUs, with HQQ used for the compressed text encoder
- Runs on Linux and Windows natively through the same CUDA / Gemlite deployment stack
- Cross-platform companion: also available as MLX 2-bit for Apple Silicon
Resources
- [White Paper](https://github.com/PrismML-Eng/Bonsai-Image-Demo/blob/main/bonsai-image-4b-whitepaper.pdf) — full benchmarks, kernels, and memory analysis
- [Demo repo](https://github.com/PrismML-Eng/Bonsai-Image-Demo) — one-command setup for Mac / Linux / Windows
- [Discord](https://discord.gg/prismml) — community + support
- Kernels: gemlite (fused low-bit GEMM) · HQQ (low-bit quantization runtime) · triton-windows (Windows path)
Model Overview
Ternary Weight Representation: 1.58-bit g128
Each ternary weight takes a value from {−1, 0, +1} with one shared FP16 scale per group of 128 weights:
w_i = scale_g * t_i, t_i in {−1, 0, +1}Ternary values carry log₂(3) ≈ 1.585 bits of information per weight. With one FP16 scale per group of 128, the effective storage is:
b_eff ≈ log2(3) + 16/128 ≈ 1.585 + 0.125 ≈ 1.71 bits/weightThis gives an idealized 9.4× reduction relative to FP16 for the ternary transformer layers. A small set of precision-sensitive supporting tensors remains in FP16, so the final Ternary Bonsai Image 4B diffusion transformer is 1.21 GB, a 6.4x reduction from the 7.75 GB FP16 FLUX.2 Klein 4B transformer.
The ternary representation is applied to the matrix-heavy transformer layers, including Q / K / V projections, output projections, MLP linears, and the double-stream add-K / Q / V linears. Supporting tensors (less than 5% of the total parameters) such as modulation streams, embedders, output norm, and output projection remain FP16 for image quality and stability.
The CUDA deployment uses a Gemlite INT2 packed format. Ternary values are stored in 2-bit slots, with the fourth code unused. The model-level Bonsai representation is 1.21 GB; the deployed CUDA pack is 1.54 GB on disk due to runtime packing and alignment overhead in the current Gemlite path.
Memory
CUDA deployment:
At runtime, the text encoder is offloaded after prompt encoding. During denoising, the repeated image-generation loop is dominated by the compact ternary diffusion transformer and active image-generation components rather than the full payload.
Peak HBM at 1024² on RTX 3080 is ~6.8 GiB end-to-end (transformer + VAE + activation memory).
Best Practices
- Sampler: FlowMatchEuler-discrete with 4 steps, guidance = 1.0, shift = 3.0. The model is designed for 4 steps; running more steps does not improve quality significantly and can introduce artifacts.
- Resolution: native 1024² is the design target. 512² works for quick previews.
- Aspect ratios: multiples of 32 are supported, including 832x1248 and 1248x832.
- Prompting: natural-language prompts. Negative prompts are not required.
- Runtime memory: the text encoder is offloaded after prompt encoding, so the denoising loop is memory-light.
Quickstart
Bonsai Studio (Linux / Windows)
The simplest path is the Bonsai Image Demo repo, which sets up the full Bonsai Studio (FastAPI backend + Next.js frontend) and selects gemlite automatically on Linux / Windows:
git clone https://github.com/PrismML-Eng/Bonsai-Image-Demo.git
cd Bonsai-Image-Demo
./setup.sh
./scripts/download_model.sh # ternary is the default
./scripts/serve.shOn Windows (PowerShell):
Set-ExecutionPolicy -Scope CurrentUser RemoteSigned # one-time
.\setup.ps1
.\scripts\download_model.ps1
.\scripts\serve.ps1Python API (backend_gpu)
For inference without the studio frontend:
from backend_gpu.server import build_pipeline
pipe = build_pipeline(model_id="prism-ml/bonsai-image-ternary-4B-gemlite-2bit")
image = pipe(
prompt="A bonsai tree in a quiet ceramic studio, soft morning light",
num_inference_steps=4,
guidance_scale=1.0,
height=1024,
width=1024,
).images[0]
image.save("bonsai.png")Throughput (CUDA / gemlite)
Warmed wall-clock per image, 4 denoising steps, guidance = 1.0, matched prompts and sampler settings.
The sub-2-bit pack keeps generation viable on commodity GPUs at 1024². The RTX 3080 10 GB reaches 4.5 s/image, while the 6 GB laptop RTX 3060 is the memory-constrained tail.
Benchmarks
Evaluated with matched generation settings across the comparison set on H100. GenEval uses the official 512x512 protocol. For HPSv3 and DPG-Bench, larger-backbone rows are evaluated at 1024x1024, while smaller-backbone rows are evaluated at their native 512x512 setting. Higher is better for all three benchmarks.
The benchmark results show the intended quality-footprint trade-off. Ternary Bonsai Image 4B is the quality-oriented variant: at 1.21 GB, it sits very close to FLUX.2 Klein 4B across GenEval, HPSv3, and DPG-Bench while reducing the diffusion transformer footprint by 6.4x. The binary companion is the footprint-oriented variant, reducing the diffusion transformer below 1 GB while still delivering strong benchmark results.
Together, the Bonsai Image variants move the quality-footprint frontier: they bring modern diffusion-transformer behavior into a memory range previously occupied by much smaller, lower-capability models.
Use Cases
- Local creative tooling: image generation directly on CUDA-equipped workstations and consumer GPUs
- Private generation: prompts and generated assets can remain in local or controlled environments
- Rapid iteration: lower local latency and no remote queue for iterative creative workflows
- Commodity-GPU serving: lower transformer footprint and reduced memory pressure for serving on NVIDIA GPUs
- Windows and Linux deployment: native paths through the same Gemlite deployment stack
- Enterprise and controlled inference: local or private environments for data residency and compliance-sensitive workflows
Limitations
- Ternary Bonsai Image 4B is not bit-identical to the FP16 FLUX.2 Klein 4B model; it is a compact ternary-weight deployment designed to deliver similar practical behavior at much smaller size.
- Image-generation quality remains prompt- and workflow-dependent. Small text, fine details, object counts, and strict compositional constraints should be evaluated for the target use case.
- Current commodity inference stacks do not yet expose fully native ternary execution as a standard hardware path. This release uses practical Gemlite low-bit GEMM kernels on CUDA.
- After the diffusion transformer is made compact, other components such as the VAE can become more visible memory bottlenecks. The runtime mitigates this with text-encoder offload and tiled VAE decoding.
Citation
@techreport{bonsaiimage4b,
title = {Bonsai Image 4B: Low-Bit Diffusion on Apple Silicon and Consumer GPUs},
author = {Prism ML},
year = {2026},
month = {May},
url = {https://prismml.com}
}Contact
For questions, feedback, or collaboration inquiries: contact@prismml.com
