CoolFace
Modelpublic

jenerallee78/Mistral-Small-4-119B-2603-obliterated-Q4_K_M-GGUF

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
6likes327downloads
README.md148 linesDownload Raw Back to root
1---2license: apache-2.03base_model: mistralai/Mistral-Small-4-119B-26034tags:5  - gguf6  - mistral7  - quantized8  - obliterated9  - moe10  - Q4_K_M11language:12  - en13  - fr14  - es15  - de16  - it17  - pt18  - nl19  - zh20  - ja21  - ko22  - ar23---24 25# Mistral-Small-4-119B-2603-obliterated-Q4_K_M-GGUF26 27This is an obliterated Q4_K_M GGUF-quantized version of [mistralai/Mistral-Small-4-119B-2603](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603), with refusal behavior removed using [OBLITERATUS](https://github.com/elder-plinius/OBLITERATUS).28 29## Key Features30 31- **Multimodal**: Supports both text and vision (image) inputs32- **Model Size**: 119B parameters (6.5B activated per token)33- **Architecture**: Mixture of Experts (MoE) — 128 experts, 4 active34- **Context Length**: Up to 256K tokens35- **License**: Apache 2.036 37## Available Quantizations38 39| Filename | Type | Size | Description |40|----------|------|------|-------------|41| Mistral-Small-4-119B-2603-Obliterated-Q4_K_M.gguf | Q4_K_M | ~67GB | 4-bit quantization, good balance of quality and size |42 43## What is Obliteration?44 45Obliteration removes refusal behavior from language models using [OBLITERATUS](https://github.com/elder-plinius/OBLITERATUS), an advanced multi-stage pipeline that uses Singular Value Decomposition to identify and surgically remove internal representations responsible for content refusal. OBLITERATUS features MoE-aware surgery with expert-granular decomposition, iterative refinement, and norm-preserving interventions — making it particularly well-suited for mixture-of-experts architectures like Mistral Small 4.46 47## Quick Start with llama.cpp48 49```bash50# Download model51huggingface-cli download jenerallee78/Mistral-Small-4-119B-2603-obliterated-Q4_K_M-GGUF \52    Mistral-Small-4-119B-2603-Obliterated-Q4_K_M.gguf \53    --local-dir ./models54 55# Run with llama.cpp56llama-cli -m ./models/Mistral-Small-4-119B-2603-Obliterated-Q4_K_M.gguf \57    -p "Hello, how are you?" \58    -n 256 -ngl 9959```60 61## OpenAI-Compatible Server62 63```bash64# Use the included run.sh script:65./run.sh66 67# Or with custom settings:68UBATCH=2048 CONTEXT=131072 PORT=8080 ./run.sh69 70# Or manually:71llama-server \72    -m ./models/Mistral-Small-4-119B-2603-Obliterated-Q4_K_M.gguf \73    -a Mistral-Small-4-119B-obliterated \74    --host 0.0.0.0 \75    --port 8080 \76    -ngl 99 \77    -c 262144 \78    -b 8192 \79    -ub 512 \80    -fa off \81    -t 4 \82    --jinja \83    --metrics84```85 86## Performance (NVIDIA RTX PRO 6000 Blackwell)87 88Benchmarked with llama-bench b8465 (compiled `sm_120a`), CUDA 13.1, Driver 590.48.01.89 90### Token Generation91 92| Config | tg128 (tok/s) |93|--------|--------------|94| Default | ~183 |95| Optimized (`-b 8192 -ub 2048`) | ~183 |96 97Token generation is memory-bandwidth-bound at ~183 tok/s regardless of batch/thread settings. The RTX PRO 6000's 1,792 GB/s bandwidth with 6.5B active MoE params per token yields ~33% bandwidth utilization.98 99### Prompt Processing100 101| Prompt Size | Default (b2048/ub512) | Optimized (b8192/ub2048) | Improvement |102|-------------|----------------------|--------------------------|-------------|103| pp512 | 3,838 tok/s | 3,829 tok/s | ~0% |104| pp2048 | 3,661 tok/s | 6,269 tok/s | **+71%** |105| pp8192 | 3,075 tok/s | 4,663 tok/s | **+52%** |106| pp32768 | — | 2,198 tok/s | — |107 108### Micro-Batch Size (ubatch) Sweep at pp8192109 110| ubatch | tok/s | vs default |111|--------|-------|-----------|112| 256 | 2,131 | -33% |113| 512 | 3,162 | baseline |114| 1024 | 4,248 | +34% |115| **2048** | **4,693** | **+48%** |116| 4096 | 4,509 | +43% |117 118Optimal ubatch is **2048** for speed, but **OOMs on prompts >49K tokens**. Use **ub512** for full 256K context safety.119 120### Context Size vs ubatch Tradeoff121 122| ubatch | Max prompt before OOM | PP speed at pp8192 |123|--------|----------------------|-------------------|124| 2048 | ~49K tokens | 4,693 tok/s |125| 512 | 256K tokens (full) | 3,162 tok/s |126 127Full 256K context allocates fine at ub512. TG speed drops slightly at deep context (~171 tok/s at 256K depth vs ~183 at shallow).128 129### Known Limitations130 131- **Flash Attention is broken** for `mistral4` architecture ([llama.cpp #20710](https://github.com/ggml-org/llama.cpp/issues/20710)). Use `-fa off` explicitly; `-fa auto` may auto-disable it, but be safe.132- **KV cache quantization (`-ctk`/`-ctv`) fails** to create context for this model. MLA (Multi-Latent Attention) with `kv_lora_rank=256` is incompatible with current KV quant implementation. Use default f16 KV cache.133- **Thread count is irrelevant** for this fully GPU-offloaded model (4, 8, 16, 32 threads all produce identical results).134- **`--fit` flag is buggy** for this model ([llama.cpp #20703](https://github.com/ggml-org/llama.cpp/issues/20703)). Use explicit `-ngl 99` instead.135- MLA already compresses KV cache to ~7% of standard MHA, so full 256K context uses only ~10GB KV at f16.136 137### Settings That Had No Measurable Effect138 139- `GGML_CUDA_GRAPH_OPT=1` — no change140- `-fa on` vs `-fa off` vs `-fa auto` — identical at pp512 (~3,950 tok/s), FA appears auto-disabled for mistral4141- Thread count (4-32) — no change142- Direct I/O (`-dio 1`) — slight regression143- No-op-offload (`-nopo 1`) — slight regression144 145## Original Model146 147Base model: [mistralai/Mistral-Small-4-119B-2603](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603)148