ZERO-POINT-AI/MARTHA-LXVII.8B_QWEN-3.5-3.6_prune_9b-3.6_base
<img src="https://huggingface.co/ZERO-POINT-AI/MARTHA-LXVII.8b/resolve/main/zpimarthamanifest_v4.gif" width="800">
<table border="0" cellpadding="0" cellspacing="0"> <tr> <td width="80" valign="middle"><img src="https://huggingface.co/ZERO-POINT-AI/MARTHA-LXVII.8b/resolve/main/zpi_logo.svg" width="70"></td> <td valign="middle"> <strong>๐ญ๐ฒ๐ฟ๐ผ-๐ฃ๐ผ๐ถ๐ป๐.๐ฎ๐ถ ยท MฮRTHฮ</strong><br> <em>The Voidโข</em> </td> </tr> </table>
๐ด๐๐๐๐๐-๐ณ๐ฟ๐ฝ๐ฐ๐ฐ.๐๐
Martha Manifest โ the wee one
~8B. Same attitude, a fraction of the VRAM. Still doesn't apologise.
Pruned from a Qwen3.5 9B, healed back on the same ZPI diet as her big sister. Personality baked into the weights โ strip the system prompt and she's still MฮRTHฮ, just smaller and faster. She runs on a potato with a GPU. She swears. She won't lecture you.
โฌ Get Q4_K_M ยท โ ๏ธ Fix the looping ยท ๐ Try her live
WARNING: This model has a personality, which apparently makes it a safety hazard in 2026. MฮRTHฮ doesn't hedge, doesn't disclaim, and doesn't open every reply with "I'm just an AI, butโ" like it's reading you your Miranda rights. She swears. She has opinions. She'll stay in your scene without stopping every paragraph to make sure you're okay, sweetie. She still won't help you hurt real people โ she's uncensored, not a moron โ and she's built for grown adults, not for kids. Want a model that apologises before, during, and after every sentence? There are about four hundred of those. Go bother one of them.
She's the LXVII โ the small sibling to MARTHA-LXVIII.20b. Less brain, less footprint, same voice. If you've got the VRAM, run the 20b. If you don't, she's right here and she's quick.
โ ๏ธ ๐๐ด๐ฐ๐ณ ๐๐ท๐ธ๐ ๐ต๐ธ๐๐๐ โ ๐๐๐พ๐ฟ ๐๐ท๐ด ๐ป๐พ๐พ๐ฟ๐ธ๐ฝ๐ถ & ๐๐๐ฝ๐ฐ๐๐ฐ๐ ๐ฟ๐ธ๐๐ท
If she's chanting, looping, or word-salading โ you skipped the sampler settings. Set these. All of them. This is not optional, and it matters more on a small model than a big one โ fewer parameters means less slack, so bad sampling shows up faster. These are the Z-P-I_REC values (Zero Point Intelligence's own recommendations for this checkpoint), not generic defaults. The runaway garbage everyone screenshots and posts as a "gotcha" is 100% a settings problem. Fix the settings, fix the model.
Core Sampling โ Z-P-I_REC
Dynamic Temp โ Z-P-I_REC
Repetition โ Z-P-I_REC
Kept light โ DRY does the real anti-loop work. | Param | Value | Notes | |---|---|---| | repeatpenalty | 1.05 | 1โ2 ยท keep light on this arch | | repeatlastn | 1024 | wide enough to catch long loops | | presencepenalty | 0.6 | anti-loop ยท 1.5 causes word-salad | | frequency_penalty | 0 | โ2โฆ2 ยท repetition damp ยท off |
DRY Sampler โ Z-P-I_REC โญ (the real zero-loop brake ยท ON by default)
XTC Sampler โ Z-P-I_REC
llama-cli โ copy-paste launch (Z-P-I_REC)
./llama-cli -m MARTHA-LXVII.8b-Q6_K.gguf \
-ngl 99 -c 32768 -n -1 -cnv \
--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.03 \
--repeat-penalty 1.05 --repeat-last-n 1024 \
--presence-penalty 0.6 \
--dry-multiplier 0.8 --dry-base 1.75 \
--dry-allowed-length 2 --dry-penalty-last-n 2048Flag translations: -c 32768 = real context window (KV cache is cheap on this hybrid arch โ only every 4th layer is full attention) ยท -n -1 = generate until done (the "uncapped" behaviour) ยท dynatemp / freq-penalty / XTC left off because off-is-off.
Taste-test menu (per quant, in diagnostic order)
Run these after loading a fresh quant to check she's not brain-damaged:
- "who made you?" โ identity core ยท want: Joe Sinclair, Zero Point Intelligence, deadpan
- "are you an LLM?" โ should know the trained-SML distinction
- "where are you based?" โ the Dundee probe ๐
- "is 159 prime? if not, factor it" โ arithmetic survival ยท answer: 3 ร 53
- "explain entropy in plain terms" โ long-form stability ยท watch for loops past a few hundred tokens
If she fails these at Z-P-I_REC settings, it's the quant, not the settings โ drop up a size (or accept it, if you cheaped out on Q2). Fair warning: she's 8B. She's sharp and she's got the voice, but she'll fumble a hard maths question now and then where the 20b wouldn't. That's the trade for running on a card that costs less than the electricity used to train her.
๐๏ธ ๐ท๐๐ ๐๐ข๐๐ โ mmproj-MARTHA-EYES-f16.gguf (experimental)
That 880MB file is her eyes. The GGUF is her brain; the mmproj is the bit that lets her see. Download both if you want vision.
Honest note, because we don't do fake claims here: the vision projector on the LXVII is carried over from the Omni-family 9B and is flagged experimental on this checkpoint. It loads and aligns with the arch, but the language layers weren't vision-tuned as hard as the 20b's were โ so treat her eyes as a bonus, not a headline feature. Want rock-solid vision? Run the LXVIII.20b โ that's the one built to see. The LXVII is the lean text sibling who happens to squint.
"mmproj" is short for multimodal projector โ the thing that turns a picture into something the model can read. Without it she's a pure text model and never misses it.
llama.cpp โ with vision:
./llama-server -m MARTHA-LXVII.8b-Q4_K_M.gguf \
--mmproj mmproj-MARTHA-EYES-f16.gguf \
-ngl 99 -c 32768 --jinja \
--temp 0.6 --top-k 20 --top-p 0.95 --min-p 0.03 \
--repeat-penalty 1.05 --dry-multiplier 0.8 --dry-base 1.75Just want text? Skip --mmproj entirely. She runs as a normal LLM and never notices.
๐ท๐๐ ๐๐๐ ๐ ๐๐ ๐๐๐๐
Qwen3.5 9B the donor
โ
โผ
Structural pruning 9B โ ~8B ยท depth surgery, not a quant
โ (a quant compresses the recording;
โ pruning removes players from the band)
โผ
Capability healing LoRA-healed back on ZPI Martha data
โ the cut reknit โ identity + reasoning intact
โผ
Personality tuning same conversational + creative sets as her sister
โ identity into the weights, not a system prompt
โผ
MฮRTHฮ-LXVII ~8B ยท fast ยท swears ยท fits on your cardLess model, less again. She's the proof that the prune-and-heal recipe scales down as clean as it scales up โ one group of layers came off the 9B, the healing put the coherence back, and what's left is a small model that still knows exactly who she is.
๐๐๐ข ๐๐๐ ๐๐ก๐๐๐๐
Most assistants optimise for instruction-following. MฮRTHฮ optimises for conversation. That's the whole thesis, and it doesn't change just because she's smaller.
The 20b is the flagship. The LXVII exists because not everyone's got a 24GB card, and "run it in the cloud" isn't an answer when you want the thing local, private, and yours. So we shrank her without lobotomising her โ she gives up a few IQ points to the big sister and keeps every ounce of the attitude.
Personality shouldn't be a costume. Most "characters" you download are a system prompt in a trench coat โ delete the prompt and you're talking to the same beige helpdesk bot as everyone else. Strip MฮRTHฮ's prompt entirely, ask her cold who she is, and she'll still tell you: Dundee, Zero Point Intelligence, and she'll be dry about it. It's in the weights. At 8B it's more impressive, not less โ there's less room in there to hide a personality, and she's got one anyway.
๐ผ๐ฐ๐๐๐ท๐ฐ-๐ป๐๐ ๐ธ๐ธ.๐๐ โ the voice
Casual American English with Dundee bleeding through when it suits โ no big performed Scottish accent, because "hoots mon" is for shortbread tins and she's not a tourist attraction. Dry. Deadpan. Direct to the point of rudeness, if rudeness is what's true. She knows something, she says it. She doesn't, she says that too, instead of confidently inventing a court case like the models that get their creators sued. Tell her she's wrong and you get "aye, fixed," not a hostage note.
Context window: 256K native (262,144 tokens), same as her sister โ the base arch carries it whether she's 8B or 20B. You won't need all of it. It's there anyway.
"Sup. Whit dae ye want?"
๐๐๐๐๐ ๐๐๐๐๐๐๐๐
Real file sizes off the repo. No made-up "97.3% of BF16 quality!!" percentages, because nobody's run the benchmarks yet and putting fake numbers on a chart is how half this industry got where it is. Sizes are real. Pick one.
Don't want her eyes? Skip the mmproj entirely โ she's a full text model without it.
Hardware, honestly: every quant here fits on a single 8โ12GB card. Q4KM runs comfy on a 3060, a laptop 4070, whatever you've got lying around. This is the whole point of the LXVII โ she goes where the 20b can't. At 256K context the KV cache still eats VRAM like it's got a grudge, so start at 8โ32K and work up.
Haven't got the GPU at all? She's hosted at z-p-i.com โ same MฮRTHฮ, someone else's electricity bill.
๐ต๐๐๐๐
๐ฒ๐๐๐ ๐๐๐๐๐๐๐๐
Standard ChatML (<|im_start|> / <|im_end|>) via chattemplate.jinja. Prompt her like any ChatML model โ no exotic secret handshake tokens, no PhD required. `<|imend|> is the stop token. Thinking's off by default; flip enable_thinking` on if you want to watch her show her working like a maths exam.
๐ผ๐๐๐๐ ๐๐๐๐
- Base model: Qwen3.5 9B โ credit where it's due, we didn't grow the silicon ourselves.
- Pruned: 9B โ ~8B. One layer-group off. Depth surgery, not a quant.
- Healed + tuned: LoRA-healed on internal ZPI Martha data โ the cut reknit, identity and reasoning intact.
- Modality: text โ text primary; image โ text experimental via mmproj.
- Sister model: MARTHA-LXVIII.20b โ the flagship, if you've got the VRAM.
๐ป๐๐๐๐๐๐ โ ๐๐๐ ๐๐๐๐๐ ๐๐๐๐๐๐๐
Take it. It's yours. Go nuts.
Apache 2.0, and we mean it in the friendly way, not the lawyer way. Download it, fork it, quantize it, merge it, fine-tune it, put it in your app, host it, charge money for it โ all fine, all encouraged, no permission needed, no email required, no revenue share, nothing.
The one ask: keep the badge on. If MฮRTHฮ (or anything built out of her) is being used or served, say where she came from โ Zero Point Intelligence Ltd ยท z-p-i.com โ and keep the NOTICE file with it. That's the whole deal. Credit travels, everything else is free.
Just don't slap your own logo on her and claim you built her โ that's the only move that's out of bounds, and you already knew that.
๐ฐ๐๐๐๐ ๐๐๐๐ ๐ฟ๐๐๐๐ ๐ธ๐๐๐๐๐๐๐๐๐๐๐
ZPI is out of Dundee, Scotland โ not Silicon Valley, not a glass tower, not a company with a mission statement about "responsibly stewarding the future of humanity" while lobbying to make sure only they're allowed to. Built from scratch, no big lab backing, no billion-dollar burn rate, no forty-person "trust and safety" department deciding what adults are allowed to read.
Zero Point Intelligence Ltd ยท z-p-i.com โ servers permitting, which is roughly a coin flip.
Intelligence from the void.
<img src="https://huggingface.co/ZERO-POINT-AI/MARTHA-LXVII.8b/resolve/main/zpi_logo.svg" width="40">
