CoolFace
Modelpublic

Abiray/Gemory-26B-A4B-GGUF

sourceHugging Faceapache-2.0updated 14d agoView on Hugging Face
1likes1.2kdownloads
Model Card

Gemory-26B-A4B - GGUF

This repository contains GGUF quantizations of UltimateIntent/Gemory-26B-A4B-GGUF.

Gemory-26B-A4B is a 26-billion parameter Mixture of Experts (MoE) model built on the Gemma architecture foundation, fine-tuned for high-fidelity creative writing, unrestricted roleplay, and complex multi-turn conversational depth. All quantizations were generated directly from the uncompressed 50.5 GB BF16 base model and sanity-tested for generation integrity.


Quantization Breakdown

FilenameQuant MethodFile SizeRecommended RAM/VRAMDescription / Quality
`Gemory-26B-A4B-Q6_K.gguf`Q6_K22.6 GB~26 GB+Near-lossless quality; closest output to base BF16 weights.
`Gemory-26B-A4B-Q5_K_M.gguf`Q5KM19.1 GB~23 GB+High quality; recommended sweet spot for 24 GB VRAM GPUs.
`Gemory-26B-A4B-Q5_K_S.gguf`Q5KS18.0 GB~22 GB+High quality with slightly faster inference and lower memory.
`Gemory-26B-A4B-Q4_K_M.gguf`Q4KM16.8 GB~20 GB+Default Recommended: Optimal balance of speed, perplexity, and memory.
`Gemory-26B-A4B-Q4_K_S.gguf`Q4KS15.5 GB~19 GB+Standard 4-bit quantization for tighter memory envelopes.
`Gemory-26B-A4B-Q3_K_M.gguf`Q3KM13.3 GB~16 GB+Low-memory 3-bit variant; preserves critical attention weights.
`Gemory-26B-A4B-Q3_K_S.gguf`Q3KS12.2 GB~15 GB+Maximum compression; fits easily on 16 GB systems.

Usage Instructions

1. llama.cpp

Command-Line Inference (`llama-cli`):

bash
./llama-cli \
  -hf Abiray/Gemory-26B-A4B-GGUF:Q4_K_M \
  -p "Write an opening scene set in a rain-soaked neon alleyway." \
  -n 512 \
  --temp 0.7