CoolFace
Modelpublic

bartowski/Rhea-72b-v0.5-GGUF-broken

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
2likes78downloads
Model Card

Llamacpp Quantizations of Rhea-72b-v0.5

Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b2440">b2440</a> for quantization.

Original model: https://huggingface.co/davidkim205/Rhea-72b-v0.5

Download a file (not the whole branch) from below:

FilenameQuant typeFile SizeDescription
Rhea-72b-v0.5-Q8_0.ggufQ8_076.82GBExtremely high quality, generally unneeded but max available quant.
Rhea-72b-v0.5-Q6_K.ggufQ6_K59.31GBVery high quality, near perfect, recommended.
Rhea-72b-v0.5-Q5_K_M.ggufQ5KM51.30GBHigh quality, very usable.
Rhea-72b-v0.5-Q5_K_S.ggufQ5KS49.88GBHigh quality, very usable.
Rhea-72b-v0.5-Q5_0.ggufQ5_049.88GBHigh quality, older format, generally not recommended.
Rhea-72b-v0.5-Q4_K_M.ggufQ4KM43.76GBGood quality, similar to 4.25 bpw.
Rhea-72b-v0.5-Q4_K_S.ggufQ4KS41.28GBSlightly lower quality with small space savings.
Rhea-72b-v0.5-IQ4_NL.ggufIQ4_NL41.25GBGood quality, similar to Q4KS, new method of quanting,
Rhea-72b-v0.5-IQ4_XS.ggufIQ4_XS39.09GBDecent quality, new method with similar performance to Q4.
Rhea-72b-v0.5-Q4_0.ggufQ4_041.00GBDecent quality, older format, generally not recommended.
Rhea-72b-v0.5-IQ3_M.ggufIQ3_M33.25GBMedium-low quality, new method with decent performance.
Rhea-72b-v0.5-IQ3_S.ggufIQ3_S31.56GBLower quality, new method with decent performance, recommended over Q3 quants.
Rhea-72b-v0.5-Q3_K_L.ggufQ3KL38.48GBLower quality but usable, good for low RAM availability.
Rhea-72b-v0.5-Q3_K_M.ggufQ3KM35.27GBEven lower quality.
Rhea-72b-v0.5-Q3_K_S.ggufQ3KS31.56GBLow quality, not recommended.
Rhea-72b-v0.5-Q2_K.ggufQ2_K27.07GBExtremely low quality, not recommended.

Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski