CoolFace
Modelpublic

jenerallee78/Mistral-Small-4-119B-2603-obliterated-Q4_K_M-GGUF

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
6likes330downloads
Model Card

Mistral-Small-4-119B-2603-obliterated-Q4KM-GGUF

This is an obliterated Q4KM GGUF-quantized version of mistralai/Mistral-Small-4-119B-2603, with refusal behavior removed using OBLITERATUS.

Key Features

  • —Multimodal: Supports both text and vision (image) inputs
  • —Model Size: 119B parameters (6.5B activated per token)
  • —Architecture: Mixture of Experts (MoE) — 128 experts, 4 active
  • —Context Length: Up to 256K tokens
  • —License: Apache 2.0

Available Quantizations

FilenameTypeSizeDescription
Mistral-Small-4-119B-2603-Obliterated-Q4KM.ggufQ4KM~67GB4-bit quantization, good balance of quality and size

What is Obliteration?

Obliteration removes refusal behavior from language models using OBLITERATUS, an advanced multi-stage pipeline that uses Singular Value Decomposition to identify and surgically remove internal representations responsible for content refusal. OBLITERATUS features MoE-aware surgery with expert-granular decomposition, iterative refinement, and norm-preserving interventions — making it particularly well-suited for mixture-of-experts architectures like Mistral Small 4.

Quick Start with llama.cpp

bash
# Download model
huggingface-cli download jenerallee78/Mistral-Small-4-119B-2603-obliterated-Q4_K_M-GGUF \
    Mistral-Small-4-119B-2603-Obliterated-Q4_K_M.gguf \
    --local-dir ./models

# Run with llama.cpp
llama-cli -m ./models/Mistral-Small-4-119B-2603-Obliterated-Q4_K_M.gguf \
    -p "Hello, how are you?" \
    -n 256 -ngl 99

OpenAI-Compatible Server

bash
# Use the included run.sh script:
./run.sh

# Or with custom settings:
UBATCH=2048 CONTEXT=131072 PORT=8080 ./run.sh

# Or manually:
llama-server \
    -m ./models/Mistral-Small-4-119B-2603-Obliterated-Q4_K_M.gguf \
    -a Mistral-Small-4-119B-obliterated \
    --host 0.0.0.0 \
    --port 8080 \
    -ngl 99 \
    -c 262144 \
    -b 8192 \
    -ub 512 \
    -fa off \
    -t 4 \
    --jinja \
    --metrics

Performance (NVIDIA RTX PRO 6000 Blackwell)

Benchmarked with llama-bench b8465 (compiled sm_120a), CUDA 13.1, Driver 590.48.01.

Token Generation

Configtg128 (tok/s)
Default~183
Optimized (-b 8192 -ub 2048)~183

Token generation is memory-bandwidth-bound at ~183 tok/s regardless of batch/thread settings. The RTX PRO 6000's 1,792 GB/s bandwidth with 6.5B active MoE params per token yields ~33% bandwidth utilization.

Prompt Processing

Prompt SizeDefault (b2048/ub512)Optimized (b8192/ub2048)Improvement
pp5123,838 tok/s3,829 tok/s~0%
pp20483,661 tok/s6,269 tok/s+71%
pp81923,075 tok/s4,663 tok/s+52%
pp32768—2,198 tok/s—

Micro-Batch Size (ubatch) Sweep at pp8192

ubatchtok/svs default
2562,131-33%
5123,162baseline
10244,248+34%
20484,693+48%
40964,509+43%

Optimal ubatch is 2048 for speed, but OOMs on prompts >49K tokens. Use ub512 for full 256K context safety.

Context Size vs ubatch Tradeoff

ubatchMax prompt before OOMPP speed at pp8192
2048~49K tokens4,693 tok/s
512256K tokens (full)3,162 tok/s

Full 256K context allocates fine at ub512. TG speed drops slightly at deep context (~171 tok/s at 256K depth vs ~183 at shallow).

Known Limitations

  • —Flash Attention is broken for mistral4 architecture (llama.cpp #20710). Use -fa off explicitly; -fa auto may auto-disable it, but be safe.
  • —KV cache quantization (`-ctk`/`-ctv`) fails to create context for this model. MLA (Multi-Latent Attention) with kv_lora_rank=256 is incompatible with current KV quant implementation. Use default f16 KV cache.
  • —Thread count is irrelevant for this fully GPU-offloaded model (4, 8, 16, 32 threads all produce identical results).
  • —`--fit` flag is buggy for this model (llama.cpp #20703). Use explicit -ngl 99 instead.
  • —MLA already compresses KV cache to ~7% of standard MHA, so full 256K context uses only ~10GB KV at f16.

Settings That Had No Measurable Effect

  • —GGML_CUDA_GRAPH_OPT=1 — no change
  • —-fa on vs -fa off vs -fa auto — identical at pp512 (~3,950 tok/s), FA appears auto-disabled for mistral4
  • —Thread count (4-32) — no change
  • —Direct I/O (-dio 1) — slight regression
  • —No-op-offload (-nopo 1) — slight regression

Original Model

Base model: mistralai/Mistral-Small-4-119B-2603