CoolFace
Modelpublic

avlp12/Krea-2-Turbo-Alis-MLX-mixed-4-8

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
4likes152downloads
Model Card

Krea 2 Turbo · Alis MLX (mixed-4/8)

Part of the Krea 2 Turbo · Alis MLX collection.

[krea/Krea-2-Turbo](https://huggingface.co/krea/Krea-2-Turbo) — a 12.9B single-stream MMDiT text-to-image model — re-implemented for Apple MLX and validated numerically faithful (cos 1.000000) to the original PyTorch, then shipped as a compact, near-lossless mixed 4/8-bit build (9.8 GB transformer) for Apple silicon.

[image]

[image]

Generated by this mixed-4/8 build (8-step Turbo, no guidance, 1024²). This is an independent, unofficial port — not an official Krea product and not endorsed by Krea.
Two builds available — pick by size vs fidelity: • **Krea-2-Turbo-Alis-MLX-8bit** — 14.2 GB, near-lossless (vel-cos 0.99994) • this repo: mixed-4/8 — 9.8 GB, the smallest near-lossless build (down_proj + endpoints @8-bit, rest @4-bit)

The weight is not the point — the verification is

Every stage of this port was cross-checked against the original PyTorch reference (`krea-ai/krea-2`) before the next stage was built — float32, fixed seed, each pipeline fed its own raw-prompt inputs.

[image]

StageMetric vs PyTorchResult
Text encoder (Qwen3-VL-4B)hidden states, 12 tapped layerscos 1.000000
Transformer (28-block DiT)velocity field (rel-L2 3e-5)cos 1.000000
VAE (Qwen-Image, decode)pixel cos vs 🤗 diffuserscos 0.9994
Full pipeline (end-to-end)pixels, identical injected noisecos 1.000000

Full end-to-end pixel cosine is 1.000000 (max|diff| 0.005, measured at 512²/8-step with identical injected noise; per-component parity is resolution-independent) — the MLX implementation is faithful to the PyTorch reference. The transformer loads krea/Krea-2-Turbo/turbo.safetensors with zero remapping (all 430 tensor names match the module tree), and the VAE reuses mflux's already-validated QwenVAE. Raw numbers: `VALIDATION_LOG.txt`. (The VAE's 0.9994 is lower than the full pipeline's 1.0 only because it was tested on a random latent — an OOD torture test; on the real latents the pipeline produces, it rounds to 1.000000.)


Architecture (what was ported)

  • —Transformer — SingleStreamDiT, 12.9B: 28 blocks × width 6144, GQA (48 query / 12 KV heads, head-dim 128), per-head QK-RMSNorm, learned sigmoid output gate, SwiGLU (16384), 3-axis RoPE. A text_fusion module collapses the 12 tapped encoder layers (2 layerwise blocks → Linear(12→1) → 2 refiner blocks). Predicts the flow-matching velocity.
  • —Text encoder — Qwen3-VL-4B-Instruct, text-only, pure-MLX. For text-only conditioning the mRoPE collapses to standard rope.
  • —VAE — AutoencoderKLQwenImage (the Qwen-Image VAE), via mflux QwenVAE.
  • —Sampler — flow-matching Euler, 8 steps, guidance 0 (distilled Turbo).

Quickstart

Requires an Apple-silicon Mac (M1+) with ≥ 24 GB unified memory (32 GB+ recommended; 16 GB will run out of memory at 1024² — use --width/--height 512). On macOS the commands are `python3` (not python). Source code, the web UI, and the full validation harness are on GitHub: [github.com/avlp12/krea2_alis_mlx](https://github.com/avlp12/krea2_alis_mlx).

🖼️ Web UI — easiest (beginners start here)

bash
python3 -m pip install mlx transformers "mflux>=0.18,<0.19" huggingface_hub gradio
hf download avlp12/Krea-2-Turbo-Alis-MLX-mixed-4-8 --local-dir krea2-mlx
cd krea2-mlx
python3 app.py            # opens http://localhost:7860 — type a prompt, click Generate ✨

The first run also downloads the Qwen3-VL-4B encoder, VAE, and tokenizer from `krea/Krea-2-Turbo` (you accept Krea's license there), so give it a few minutes; only the mixed-4/8 transformer lives in this repo. A 1024×1024 image takes ~50 s on an M3 Ultra (8 steps; slower chips take longer). An NSFW safety filter runs by default (redacts explicit outputs; disable with the UI toggle, --no-safety, or KREA2_DISABLE_SAFETY=1).

💡 The UI's Model dropdown switches between mixed-4/8 and 8-bit — the other build downloads on first use. (CLI: add --precision 8bit.)

⌨️ Command line

bash
python3 generate.py "a red fox in the snow, photorealistic" --out fox.png

Flags: --width/--height 512|768|1024, --steps 8, --seed 0, --num-images 2.

🐍 Python

python
from krea2.pipeline import Krea2Pipeline

pipe = Krea2Pipeline("transformer_mixed_4_8.safetensors", precision="mixed-4-8")
img = pipe.generate("a neon city street at night in the rain", width=1024, height=1024,
                    steps=8, seed=0)[0]
img.save("out.png")

Run full precision instead (pulls turbo.safetensors from Krea): python3 generate.py "…" --precision bf16.


Quantization — and an honest note

Sensitivity-graded mixed 4/8-bit on the 28-block bulk (224 matmuls), group-size 64: 8-bit for down_proj and the first/last-2 blocks' attention (the most quantization-sensitive spots); 4-bit for the rest of the attention + SwiGLU. Everything precision-critical stays bf16 — first / last, time-embeddings, the whole text_fusion module (incl. the Linear(12→1) projector), and all norms / modulation. The text encoder and VAE are bf16.

BuildSize (transformer)Velocity cos vs bf16 (mean / min)Per-step latency @1024²
bf16 (reference)25.6 GB—~5800 ms
8-bit (sibling repo)14.2 GB0.99994 / 0.99959~5990 ms
mixed-4/8 · this release9.8 GB0.99824 / 0.98710~5990 ms
4-bit8.2 GB0.99760 / 0.98666~5990 ms

Velocity cosine = mean / worst-case-min over 12 prompts × 8 denoising steps, on a fixed bf16 trajectory. Raw output: [`VALIDATION_LOG.txt`](VALIDATION_LOG.txt).

Why mixed-4/8? It's the smallest near-lossless build — only 1.6 GB more than plain 4-bit but measurably better on average (mean cos 0.99824 vs 0.99760, and a slightly higher worst-case step 0.98710 vs 0.98666), because down_proj and the endpoint blocks (which set/read out the few-step trajectory) keep 8 bits. MXFP4/MXFP8 were also tested and rejected — at comparable size, MLX's affine quant reconstructs these weights better.

Latency is identical across bit-widths — generation is attention-bound (attention isn't quantized), so quantization gives no speedup here; it only shrinks the download. On a big-RAM Mac, prefer bf16 or the 8-bit build. Quality screened by per-step velocity cosine on a fixed trajectory (not final-pixel cosine — an 8-step ODE makes that conflate benign trajectory divergence with real degradation), across multiple prompts; recipe 3-lens reviewed (codex + red-team + blank-slate).


License & attribution

This is a modified derivative of `krea/Krea-2-Turbo`, distributed under the [Krea 2 Community License](https://krea.ai/krea-2-licensing) (a copy is in `LICENSE`; attribution in `NOTICE`). By using these weights you agree to that license. In particular:

  • —Naming. Per the license, derivative model names begin with "Krea".
  • —Commercial use is permitted only if your total annual revenue is under $1,000,000 USD; otherwise you need an enterprise license from Krea (opensource@krea.ai).
  • —Content filtering. You must implement reasonable content-filtering safeguards in any deployment, and disclose AI-generated content where required by law. (The app ships a built-in pure-MLX NSFW filter — no PyTorch needed — on by default; if you disable it or deploy publicly the obligation is yours.)
  • —Acceptable Use Policy. Your use must comply with Krea's Acceptable Use Policy, which the license incorporates by reference.
  • —Not endorsed. Independent port; not an official Krea product and not endorsed by Krea.

You own the images you generate (subject to the license); Krea claims no ownership of outputs.

Credits

  • —Krea.ai — Krea 2 (base model & reference code)
  • —Qwen — Qwen-Image VAE & Qwen3-VL-4B text encoder
  • —[mflux](https://github.com/filipstrand/mflux) — MLX diffusion framework (VAE reused from here)

📦 Source code, web UI & validation harness: github.com/avlp12/krea2_alis_mlx

Part of the Alis MLX line — see also `avlp12/Krea-2-Turbo-Alis-MLX-8bit`, `avlp12/Lance-3B-Alis-MLX-Traced`, `avlp12/GLM-5.2-Alis-MLX-Dynamic-3.5bpw`.