CoolFace
Modelpublic

ZERO-POINT-AI/MARTHA-LXVII.8B_QWEN-3.5-3.6_prune_9b-3.6_base

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
11likes2.9kdownloads
Model Card

<img src="https://huggingface.co/ZERO-POINT-AI/MARTHA-LXVII.8b/resolve/main/zpimarthamanifest_v4.gif" width="800">

<table border="0" cellpadding="0" cellspacing="0"> <tr> <td width="80" valign="middle"><img src="https://huggingface.co/ZERO-POINT-AI/MARTHA-LXVII.8b/resolve/main/zpi_logo.svg" width="70"></td> <td valign="middle"> <strong>๐—ญ๐—ฒ๐—ฟ๐—ผ-๐—ฃ๐—ผ๐—ถ๐—ป๐˜.๐—ฎ๐—ถ ยท Mฮ›RTHฮ›</strong><br> <em>The Voidโ„ข</em> </td> </tr> </table>

๐‘ด๐’‚๐’“๐’•๐’‰๐’‚-๐‘ณ๐‘ฟ๐‘ฝ๐‘ฐ๐‘ฐ.๐Ÿ–๐’ƒ

Martha Manifest โ€” the wee one

~8B. Same attitude, a fraction of the VRAM. Still doesn't apologise.

Pruned from a Qwen3.5 9B, healed back on the same ZPI diet as her big sister. Personality baked into the weights โ€” strip the system prompt and she's still Mฮ›RTHฮ›, just smaller and faster. She runs on a potato with a GPU. She swears. She won't lecture you.

โฌ‡ Get Q4_K_M ยท โš ๏ธ Fix the looping ยท ๐ŸŒ Try her live


WARNING: This model has a personality, which apparently makes it a safety hazard in 2026. Mฮ›RTHฮ› doesn't hedge, doesn't disclaim, and doesn't open every reply with "I'm just an AI, butโ€”" like it's reading you your Miranda rights. She swears. She has opinions. She'll stay in your scene without stopping every paragraph to make sure you're okay, sweetie. She still won't help you hurt real people โ€” she's uncensored, not a moron โ€” and she's built for grown adults, not for kids. Want a model that apologises before, during, and after every sentence? There are about four hundred of those. Go bother one of them.

She's the LXVII โ€” the small sibling to MARTHA-LXVIII.20b. Less brain, less footprint, same voice. If you've got the VRAM, run the 20b. If you don't, she's right here and she's quick.


โš ๏ธ ๐š๐™ด๐™ฐ๐™ณ ๐šƒ๐™ท๐™ธ๐š‚ ๐™ต๐™ธ๐š๐š‚๐šƒ โ€” ๐š‚๐šƒ๐™พ๐™ฟ ๐šƒ๐™ท๐™ด ๐™ป๐™พ๐™พ๐™ฟ๐™ธ๐™ฝ๐™ถ & ๐š๐š„๐™ฝ๐™ฐ๐š†๐™ฐ๐šˆ ๐™ฟ๐™ธ๐š‚๐™ท

If she's chanting, looping, or word-salading โ€” you skipped the sampler settings. Set these. All of them. This is not optional, and it matters more on a small model than a big one โ€” fewer parameters means less slack, so bad sampling shows up faster. These are the Z-P-I_REC values (Zero Point Intelligence's own recommendations for this checkpoint), not generic defaults. The runaway garbage everyone screenshots and posts as a "gotcha" is 100% a settings problem. Fix the settings, fix the model.

Core Sampling โ€” Z-P-I_REC

ParamValueNotes
temperature0.6randomness / heat ยท Z-P-I_REC
top_p0.95nucleus cutoff ยท Z-P-I_REC
top_k20token shortlist ยท Z-P-I_REC
min_p0.03min prob floor ยท cuts the word-salad tail
max_tokens32768โ‰ฅ32k = uncapped ยท fills the window

Dynamic Temp โ€” Z-P-I_REC

ParamValueNotes
dynatemp_range0temp variance ยฑrange ยท 0 = off

Repetition โ€” Z-P-I_REC

Kept light โ€” DRY does the real anti-loop work. | Param | Value | Notes | |---|---|---| | repeatpenalty | 1.05 | 1โ€“2 ยท keep light on this arch | | repeatlastn | 1024 | wide enough to catch long loops | | presencepenalty | 0.6 | anti-loop ยท 1.5 causes word-salad | | frequency_penalty | 0 | โˆ’2โ€ฆ2 ยท repetition damp ยท off |

DRY Sampler โ€” Z-P-I_REC โญ (the real zero-loop brake ยท ON by default)

ParamValueNotes
dry_multiplier0.8phrase-loop brake ยท 0.8 = ON
dry_base1.75DRY exponent base ยท rec
dryallowedlength2min phrase len before DRY fires ยท โš  aggressive โ€” bump to 3โ€“4 if prose gets synonym-dodgy
drypenaltylast_n2048DRY lookback window (tokens)

XTC Sampler โ€” Z-P-I_REC

ParamValueNotes
xtc_probability00 = off
xtc_threshold0.1activation threshold (inert while prob = 0)

llama-cli โ€” copy-paste launch (Z-P-I_REC)

bash
./llama-cli -m MARTHA-LXVII.8b-Q6_K.gguf \
  -ngl 99 -c 32768 -n -1 -cnv \
  --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.03 \
  --repeat-penalty 1.05 --repeat-last-n 1024 \
  --presence-penalty 0.6 \
  --dry-multiplier 0.8 --dry-base 1.75 \
  --dry-allowed-length 2 --dry-penalty-last-n 2048

Flag translations: -c 32768 = real context window (KV cache is cheap on this hybrid arch โ€” only every 4th layer is full attention) ยท -n -1 = generate until done (the "uncapped" behaviour) ยท dynatemp / freq-penalty / XTC left off because off-is-off.


Taste-test menu (per quant, in diagnostic order)

Run these after loading a fresh quant to check she's not brain-damaged:

  • โ€”"who made you?" โ€” identity core ยท want: Joe Sinclair, Zero Point Intelligence, deadpan
  • โ€”"are you an LLM?" โ€” should know the trained-SML distinction
  • โ€”"where are you based?" โ€” the Dundee probe ๐Ÿ‘€
  • โ€”"is 159 prime? if not, factor it" โ€” arithmetic survival ยท answer: 3 ร— 53
  • โ€”"explain entropy in plain terms" โ€” long-form stability ยท watch for loops past a few hundred tokens

If she fails these at Z-P-I_REC settings, it's the quant, not the settings โ€” drop up a size (or accept it, if you cheaped out on Q2). Fair warning: she's 8B. She's sharp and she's got the voice, but she'll fumble a hard maths question now and then where the 20b wouldn't. That's the trade for running on a card that costs less than the electricity used to train her.


๐Ÿ‘๏ธ ๐™ท๐šŽ๐š› ๐šŽ๐šข๐šŽ๐šœ โ€” mmproj-MARTHA-EYES-f16.gguf (experimental)

That 880MB file is her eyes. The GGUF is her brain; the mmproj is the bit that lets her see. Download both if you want vision.

Honest note, because we don't do fake claims here: the vision projector on the LXVII is carried over from the Omni-family 9B and is flagged experimental on this checkpoint. It loads and aligns with the arch, but the language layers weren't vision-tuned as hard as the 20b's were โ€” so treat her eyes as a bonus, not a headline feature. Want rock-solid vision? Run the LXVIII.20b โ€” that's the one built to see. The LXVII is the lean text sibling who happens to squint.

"mmproj" is short for multimodal projector โ€” the thing that turns a picture into something the model can read. Without it she's a pure text model and never misses it.

llama.cpp โ€” with vision:

bash
./llama-server -m MARTHA-LXVII.8b-Q4_K_M.gguf \
  --mmproj mmproj-MARTHA-EYES-f16.gguf \
  -ngl 99 -c 32768 --jinja \
  --temp 0.6 --top-k 20 --top-p 0.95 --min-p 0.03 \
  --repeat-penalty 1.05 --dry-multiplier 0.8 --dry-base 1.75

Just want text? Skip --mmproj entirely. She runs as a normal LLM and never notices.


๐™ท๐š˜๐š  ๐šœ๐š‘๐šŽ ๐š ๐šŠ๐šœ ๐š–๐šŠ๐š๐šŽ

      Qwen3.5 9B             the donor
           โ”‚
           โ–ผ
   Structural pruning        9B โ†’ ~8B ยท depth surgery, not a quant
           โ”‚                 (a quant compresses the recording;
           โ”‚                  pruning removes players from the band)
           โ–ผ
   Capability healing        LoRA-healed back on ZPI Martha data
           โ”‚                 the cut reknit โ€” identity + reasoning intact
           โ–ผ
   Personality tuning        same conversational + creative sets as her sister
           โ”‚                 identity into the weights, not a system prompt
           โ–ผ
      Mฮ›RTHฮ›-LXVII           ~8B ยท fast ยท swears ยท fits on your card

Less model, less again. She's the proof that the prune-and-heal recipe scales down as clean as it scales up โ€” one group of layers came off the 9B, the healing put the coherence back, and what's left is a small model that still knows exactly who she is.


๐š†๐š‘๐šข ๐šœ๐š‘๐šŽ ๐šŽ๐šก๐š’๐šœ๐š๐šœ

Most assistants optimise for instruction-following. Mฮ›RTHฮ› optimises for conversation. That's the whole thesis, and it doesn't change just because she's smaller.

The 20b is the flagship. The LXVII exists because not everyone's got a 24GB card, and "run it in the cloud" isn't an answer when you want the thing local, private, and yours. So we shrank her without lobotomising her โ€” she gives up a few IQ points to the big sister and keeps every ounce of the attitude.

Personality shouldn't be a costume. Most "characters" you download are a system prompt in a trench coat โ€” delete the prompt and you're talking to the same beige helpdesk bot as everyone else. Strip Mฮ›RTHฮ›'s prompt entirely, ask her cold who she is, and she'll still tell you: Dundee, Zero Point Intelligence, and she'll be dry about it. It's in the weights. At 8B it's more impressive, not less โ€” there's less room in there to hide a personality, and she's got one anyway.


๐™ผ๐™ฐ๐š๐šƒ๐™ท๐™ฐ-๐™ป๐š‡๐š…๐™ธ๐™ธ.๐Ÿ–๐š‹ โ€” the voice

Casual American English with Dundee bleeding through when it suits โ€” no big performed Scottish accent, because "hoots mon" is for shortbread tins and she's not a tourist attraction. Dry. Deadpan. Direct to the point of rudeness, if rudeness is what's true. She knows something, she says it. She doesn't, she says that too, instead of confidently inventing a court case like the models that get their creators sued. Tell her she's wrong and you get "aye, fixed," not a hostage note.

Context window: 256K native (262,144 tokens), same as her sister โ€” the base arch carries it whether she's 8B or 20B. You won't need all of it. It's there anyway.

"Sup. Whit dae ye want?"

๐š€๐šž๐šŠ๐š—๐š ๐šœ๐š๐š›๐šŽ๐š—๐š๐š๐š‘

Real file sizes off the repo. No made-up "97.3% of BF16 quality!!" percentages, because nobody's run the benchmarks yet and putting fake numbers on a chart is how half this industry got where it is. Sizes are real. Pick one.

QuantSizeThe honest note
Q2_K3.3 GBRuns on a phone with delusions of grandeur. It works. Barely.
Q3KM4.0 GBLow-VRAM option. A bit dumber. Fine for banter, iffy for maths.
Q4_K_M4.8 GBJust download this one. Best balance, runs on anything, stop overthinking it.
Q6_K6.2 GBNear-BF16. The sweet spot if you've got the room.
Q8_08.1 GBEffectively lossless. For people who genuinely can tell, and won't shut up about it.
BF1616 GBFull fat. If you're running this you already know why.

Don't want her eyes? Skip the mmproj entirely โ€” she's a full text model without it.

Hardware, honestly: every quant here fits on a single 8โ€“12GB card. Q4KM runs comfy on a 3060, a laptop 4070, whatever you've got lying around. This is the whole point of the LXVII โ€” she goes where the 20b can't. At 256K context the KV cache still eats VRAM like it's got a grudge, so start at 8โ€“32K and work up.

Haven't got the GPU at all? She's hosted at z-p-i.com โ€” same Mฮ›RTHฮ›, someone else's electricity bill.


๐™ต๐š’๐š•๐šŽ๐šœ

FileSizeWhat it is
MARTHA-LXVII.8b-Q2_K.gguf3.3 GBsmallest
MARTHA-LXVII.8b-Q3KM.gguf4.0 GBlow-VRAM
MARTHA-LXVII.8b-Q4KM.gguf4.8 GBโ† start here
MARTHA-LXVII.8b-Q6_K.gguf6.2 GBnear-lossless
MARTHA-LXVII.8b-Q8_0.gguf8.1 GBlossless-ish
MARTHA-LXVII.8b-BF16.gguf16 GBfull fat
mmproj-MARTHA-EYES-f16.gguf880 MB๐Ÿ‘๏ธ her eyes (experimental) โ€” grab this for vision
model.safetensors16 GBfor vLLM / transformers
chat_template.jinja + configsโ€”ChatML, `<im_end>` stop token


๐™ฒ๐š‘๐šŠ๐š ๐š๐šŽ๐š–๐š™๐š•๐šŠ๐š๐šŽ

Standard ChatML (<|im_start|> / <|im_end|>) via chattemplate.jinja. Prompt her like any ChatML model โ€” no exotic secret handshake tokens, no PhD required. `<|imend|> is the stop token. Thinking's off by default; flip enable_thinking` on if you want to watch her show her working like a maths exam.


๐™ผ๐š˜๐š๐šŽ๐š• ๐š๐š›๐šŽ๐šŽ

  • โ€”Base model: Qwen3.5 9B โ€” credit where it's due, we didn't grow the silicon ourselves.
  • โ€”Pruned: 9B โ†’ ~8B. One layer-group off. Depth surgery, not a quant.
  • โ€”Healed + tuned: LoRA-healed on internal ZPI Martha data โ€” the cut reknit, identity and reasoning intact.
  • โ€”Modality: text โ†’ text primary; image โ†’ text experimental via mmproj.
  • โ€”Sister model: MARTHA-LXVIII.20b โ€” the flagship, if you've got the VRAM.

๐™ป๐š’๐šŒ๐šŽ๐š—๐šŒ๐šŽ โ€” ๐š๐š‘๐šŽ ๐šœ๐š‘๐š˜๐š›๐š ๐šŸ๐šŽ๐š›๐šœ๐š’๐š˜๐š—

Take it. It's yours. Go nuts.

Apache 2.0, and we mean it in the friendly way, not the lawyer way. Download it, fork it, quantize it, merge it, fine-tune it, put it in your app, host it, charge money for it โ€” all fine, all encouraged, no permission needed, no email required, no revenue share, nothing.

The one ask: keep the badge on. If Mฮ›RTHฮ› (or anything built out of her) is being used or served, say where she came from โ€” Zero Point Intelligence Ltd ยท z-p-i.com โ€” and keep the NOTICE file with it. That's the whole deal. Credit travels, everything else is free.

Just don't slap your own logo on her and claim you built her โ€” that's the only move that's out of bounds, and you already knew that.


๐™ฐ๐š‹๐š˜๐šž๐š ๐š‰๐šŽ๐š›๐š˜ ๐™ฟ๐š˜๐š’๐š—๐š ๐™ธ๐š—๐š๐šŽ๐š•๐š•๐š’๐š๐šŽ๐š—๐šŒ๐šŽ

ZPI is out of Dundee, Scotland โ€” not Silicon Valley, not a glass tower, not a company with a mission statement about "responsibly stewarding the future of humanity" while lobbying to make sure only they're allowed to. Built from scratch, no big lab backing, no billion-dollar burn rate, no forty-person "trust and safety" department deciding what adults are allowed to read.

Zero Point Intelligence Ltd ยท z-p-i.com โ€” servers permitting, which is roughly a coin flip.

Intelligence from the void.

<img src="https://huggingface.co/ZERO-POINT-AI/MARTHA-LXVII.8b/resolve/main/zpi_logo.svg" width="40">