CoolFace
Modelpublic

Eviation/flux-imatrix

sourceHugging Faceotherupdated 1y agoView on Hugging Face
4likes4.1kdownloads
Model Card

Supported?

Expect broken or faulty items for the time being. Use at your own discretion.

  • ComfyUI-GGUF: all? (CPU/CUDA)
  • Fast dequant: BF16, Q80, Q51, Q50, Q41, Q40, Q6K, Q5K, Q4K, Q3K, Q2K
  • Slow dequant: others via GGUF/NumPy
  • Forge: TBC
  • stable-diffusion.cpp: llama.cpp Feature-matrix
  • CPU: all
  • Cuda: all?
  • Vulkan: >= Q3KS, > IQ4S; [PR IQ1S, IQ1M](https://github.com/ggerganov/llama.cpp/pull/11528) [PR IQ4XS](https://github.com/ggerganov/llama.cpp/pull/11501)
  • other: ?

Disco

Dynamic quantization:

  • timein.inlayer: Q80/Q6K
  • finallayer, vectorin.inlayer, guidancein: Q8_0
  • vectorin.outlayer, timein.outlayer, txtin, imgin: F16
  • single_blocks.[> 10 && < 37].modulation.lin: one down?
FilenameQuant typeFile SizeDescription / L2 Loss Step 25Example Image

Caesar

Combined imatrix multiple images 512x512 and 768x768, 25, 30 and 50 steps city96/flux1-dev-Q8_0 euler

data: load_imatrix: loaded 314 importance matrix entries from imatrix_caesar.dat computed on 475 chunks

Using llama.cpp quantize cae9fb4 with modified lcpp.patch.

Dynamic quantization:

  • imgin, guidancein.inlayer, finallayer.linear: f32/bf16/f16
  • guidancein, finallayer: bf16/f16
  • img_attn.qkv, linear1: some layers two bits up
  • txtmod.lin, txtmlp, txt_attn.proj: some layers one bit down

Experimental from f16

FilenameQuant typeFile SizeDescription / L2 Loss Step 25Example Image
flux1-dev-IQ1_S.ggufIQ1_S2.41GBworst / 173Example
flux1-dev-TQ1_0.ggufTQ1_02.64GBworst / 195Example
flux1-dev-IQ1_M.ggufIQ1_M2.72GBworst / 171Example
flux1-dev-IQ2_XXS.ggufIQ2_XXS3.10GBworst * / 126Example
flux1-dev-TQ2_0.ggufTQ2_03.12GBworst / 202Example
flux1-dev-IQ2_XS.ggufIQ2_XS3.48GBworst / 140Example
flux1-dev-IQ2_S.ggufIQ2_S3.51GBworst / 142Example
flux1-dev-IQ2_M.ggufIQ2_M3.84GBbad / 120Example
flux1-dev-Q2_K_S.ggufQ2KS4.00GBok * / 52Example
flux1-dev-Q2_K.ggufQ2_K4.03GBok / 55Example
flux1-dev-IQ3_XXS.ggufIQ3_XXS4.56GBok / 92Example
flux1-dev-IQ3_XS.ggufIQ3_XS5.05GBbad / 125Example
flux1-dev-Q3_K_S.ggufQ3KS5.10GBok / 48Example
flux1-dev-IQ3_S.ggufIQ3_S5.11GBbad / 123Example
flux1-dev-Q3_K_M.ggufQ3KM5.13GBok / 50Example
flux1-dev-IQ3_M.ggufIQ3_M5.14GBbad / 123Example
flux1-dev-Q3_K_L.ggufQ3KL5.17GBok / 61Example
flux1-dev-IQ4_XS.ggufIQ4_XS6.33GBgood / 33Example
flux1-dev-Q4_K_S.ggufQ4KS6.66GBgood / 22Example
flux1-dev-Q4_K_M.ggufQ4KM6.69GBgood / 21Example
flux1-dev-IQ4_NL.ggufIQ4_NL6.69GBgood / 24Example
flux1-dev-Q4_0.ggufQ4_06.81GBgood / 30Example
flux1-dev-Q4_1.ggufQ4_17.55GBgood / 27Example
flux1-dev-Q5_K_S.ggufQ5KS8.26GBnice / 21Example
flux1-dev-Q5_0.ggufQ5_08.27GBgood / 30Example
flux1-dev-Q5_K_M.ggufQ5KM8.30GBnice / 23Example
flux1-dev-Q5_1.ggufQ5_18.99GBnice * / 14Example
flux1-dev-Q6_K.ggufQ6_K9.80GBnice / 20Example
flux1-dev-Q8_0.ggufQ8_012.3GBnear perfect * / 8Example
-F1623.8GBreferenceExample
FilenameBits img_attn.qkv & linear1
flux1-dev-IQ1_S.gguf333M MMMM M111 ... 11MM MM11
flux1-dev-TQ1_0.gguf3332 2222 2111 ... 1122 2211
flux1-dev-IQ1_M.gguf3332 2222 2111 ... 1122 2211
flux1-dev-IQ2_XXS.gguf4433 3333 3222 ... 2222
flux1-dev-TQ2_0.gguf3332 2222 2111 ... 1122 2211
flux1-dev-IQ2_XS.gguf4443 3333 3222 ... 2233 3322
flux1-dev-IQ2_S.gguf4444 4444 4444 4444 4433 3222 ... 2233 3322
flux1-dev-IQ2_M.gguf4444 4444 4444 4444 4433 3222 ... 2223 3333 3322
flux1-dev-Q2_K_S.gguf4443 3333 3222 ... 2222
flux1-dev-Q2_K.gguf4443 3333 3222 ... 2233 3322
flux1-dev-IQ3_XXS.gguf444S SSSS S333 ... 3333
flux1-dev-IQ3_XS.gguf444S SSSS S333 ... 33SS SS33
flux1-dev-Q3_K_S.gguf5554 4444 4333 ... 3333
flux1-dev-IQ3_S.gguf5554 4444 4333 ... 3344 4433
flux1-dev-Q3_K_M.gguf5554 4444 4333 ... 3344 4433
flux1-dev-IQ3_M.gguf5554 4444 4444 4444 4433 ... 3344 4433
flux1-dev-Q3_K_L.gguf5554 4444 4444 4444 4433 ... 3344 4433
flux1-dev-IQ4_XS.gguf8885 5555 5444 ... 4444
flux1-dev-Q4_K_S.gguf8885 5555 5444 ... 4444
flux1-dev-Q4_K_M.gguf8885 5555 5555 5555 5544 ... 4444
flux1-dev-IQ4_NL.gguf8885 5555 5555 5555 5544 ... 4444
flux1-dev-Q4_0.gguf8885 5555 5444 ... 4444
flux1-dev-Q4_1.gguf8885 5555 5444 ... 4444
flux1-dev-Q5_K_S.ggufFFF6 6666 6666 6666 6655 ... 5555
flux1-dev-Q5_0.ggufFFF8 8888 8555 ... 5555
flux1-dev-Q5_K_M.ggufFFF8 8888 8666 6666 6655 ... 5555
flux1-dev-Q5_1.ggufFFF8 8888 8555 ... 5555
flux1-dev-Q6_K.ggufFFF8 8888 8666 .. 6666
flux1-dev-Q8_0.ggufFFF8 8888 .. 8888

Observations

  • More imatrix data doesn't necessarily result in better quants
  • I-quants worse than same bits k-quants?
  • Quant-dequant loss

Bravo

Combined imatrix multiple images 512x512 25 and 50 steps city96/flux1-dev-Q8_0 euler

Using llama.cpp quantize cae9fb4 with modified lcpp.patch.

Experimental from f16

FilenameQuant typeFile SizeDescription / L2 Loss Step 25Example Image
flux1-dev-IQ1_S.ggufIQ1_S2.45GBworst / 156Example
flux1-dev-IQ1_M.ggufIQ1_M2.72GBworst / 141Example
flux1-dev-IQ2_XXS.ggufIQ2_XXS3.19GBworst / 131Example
flux1-dev-IQ2_XS.ggufIQ2_XS3.56GBworst / 125-
flux1-dev-IQ2_S.ggufIQ2_S3.56GBworst / 125-
flux1-dev-IQ2_M.ggufIQ2_M3.93GBworst / 120Example
flux1-dev-Q2_K_S.ggufQ2KS4.02GBok / 56Example
flux1-dev-IQ3_XXS.ggufIQ3_XXS4.66GBTBC / 68Example
flux1-dev-IQ3_XS.ggufIQ3_XS5.22GBworse than IQ3_XXS / 115Example
flux1-dev-IQ3_S.ggufIQ3_STBCTBC-
flux1-dev-IQ3_M.ggufIQ3_MTBCTBC-
flux1-dev-Q3_K_S.ggufQ3KS5.22GBTBC / 34Example
flux1-dev-IQ4_XS.ggufIQ4_XS6.42GBTBC / 25-
flux1-dev-Q4_0.ggufQ4_06.79GBTBC / 31-
flux1-dev-IQ4_NL.ggufIQ4_NL6.79GBTBC / 21Example
flux1-dev-Q4_K_S.ggufQ4KS6.79GBTBC / 29Example
flux1-dev-Q4_1.ggufQ4_17.53GBTBC / 24-
flux1-dev-Q5_0.ggufQ5_08.27GBTBC / 25-
flux1-dev-Q5_1.ggufQ5_1TBCTBC / 24-
flux1-dev-Q5_K_S.ggufQ5KS8.27GBTBC / 20Example
flux1-dev-Q6_K.ggufQ6_K9.84GBTBC / 19Example
flux1-dev-Q8_0.ggufQ8_0-TBC / 10-
-F1623.8GBreferenceExample

Observations

Alpha

Simple imatrix: 512x512 single image 8/20 steps city96/flux1-dev-Q3_K_S euler

data: load_imatrix: loaded 314 importance matrix entries from imatrix.dat computed on 7 chunks.

Using llama.cpp quantize cae9fb4 with modified lcpp.patch.

Experimental from q8

FilenameQuant typeFile SizeDescription / L2 Loss Step 25Example Image
flux1-dev-IQ1_S.ggufIQ1_S2.45GBworst / 152Example
-IQ1_M-broken-
flux1-dev-TQ1_0.ggufTQ1_02.63GBTBC / 220-
flux1-dev-TQ2_0.ggufTQ2_03.19GBTBC / 220-
flux1-dev-IQ2_XXS.ggufIQ2_XXS3.19GBworst / 130Example
flux1-dev-IQ2_XS.ggufIQ2_XS3.56GBworst / 129Example
flux1-dev-IQ2_S.ggufIQ2_S3.56GBworst / 129-
flux1-dev-IQ2_M.ggufIQ2_M3.93GBworst / 121-
flux1-dev-Q2_K.ggufQ2_K4.02GBTBC / 77-
flux1-dev-Q2_K_S.ggufQ2KS4.02GBok / 77Example
flux1-dev-IQ3_XXS.ggufIQ3_XXS4.66GBTBC / 130Example
flux1-dev-IQ3_XS.ggufIQ3_XS5.22GBTBC / 114-
flux1-dev-IQ3_S.ggufIQ3_S5.22GBTBC / 114-
flux1-dev-IQ3_M.ggufIQ3_M5.22GBTBC / 114-
flux1-dev-Q3_K_S.ggufQ3KS5.22GBTBC / 36Example
flux1-dev-Q3_K_M.ggufQ3KM5.36GBTBC / 42-
flux1-dev-Q3_K_L.ggufQ3KL5.36GBTBC / 42-
flux1-dev-IQ4_XS.ggufIQ4_XS6.42GBTBC / 30Example
flux1-dev-IQ4_NL.ggufIQ4_NL6.79GBTBC / 23Example
flux1-dev-Q4_0.ggufQ4_06.79GBTBC / 27-
-Q4_KTBCTBC / 27-
flux1-dev-Q4_K_S.ggufQ4KS6.79GBTBC / 26Example
flux1-dev-Q4_K_M.ggufQ4KM6.93GBTBC / 27-
flux1-dev-Q4_1.ggufQ4_17.53GBTBC / 23-
flux1-dev-Q5_K_S.ggufQ5KS8.27GBTBC / 19Example
flux1-dev-Q5_K.ggufQ5_K8.41GBTBC / 20-
-Q5KMTBCTBC-
flux1-dev-Q6_K.ggufQ6_K9.84GBTBC / 22-
-Q8_012.7GBnear perfect / 10Example
-F1623.8GBreferenceExample

Observations

Sub-quants not diferentiated as expected: IQ2XS == IQ2S, IQ3XS == IQ3S == IQ3M, Q3KM == Q3K_L.

  • Check if lcpp_sd3.patch includes more specific quant level logic
  • Extrapolate the existing level logic