CoolFace
Modelpublic

axiomofmind/Hornybot-Mara-POV

sourceHugging Faceupdated 12d agoView on Hugging Face
2likes719downloads
Model Card

Hornybot Mara POV

A fine-tune of Qwen/Qwen3.5-9B for fictional adult roleplay as Mara, a playful 28-year-old character. This POV edition narrates Mara's dialogue and actions in first person using I/me/my.

The compact system prompt is required for this behavior and is embedded in chat_template.jinja and both GGUF files. Leave the client's system field empty to use it automatically.

Developed by A Hole AI.

Files

FileFormatSizePurpose
Transformers model filesBF1618.82 GBMerged weights
Hornybot-Mara-POV-BF16.ggufBF16 GGUF17.92 GBUnquantized GGUF
Hornybot-Mara-POV-Q6_K.ggufQ6_K GGUF7.36 GBCompact local download

llama.cpp

Use a build with Qwen3.5 support. After downloading the Q6_K file:

bash
llama-server -m Hornybot-Mara-POV-Q6_K.gguf --ctx-size 32768 --flash-attn on --n-gpu-layers all --reasoning off --jinja --ui

Open http://127.0.0.1:8080 after the server starts.

SettingValue
System promptLeave empty; required default is embedded
ReasoningOff
Temperature0.7
Top-p0.9
Top-k20
Min-p0
Repetition penalty1.0
Maximum new tokens256

The chat template supplies Mara's first-person character default. A client system message is appended as extra scene context.

Transformers

python
import torch
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration

model_id = "axiomofmind/Hornybot-Mara-POV"
processor = AutoProcessor.from_pretrained(model_id)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
    model_id, dtype=torch.bfloat16, device_map="auto"
)

messages = [{"role": "user", "content": "You made it. How was your night?"}]
prompt = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
inputs = processor(text=[prompt], return_tensors="pt").to(model.device)

with torch.inference_mode():
    output = model.generate(
        **inputs, do_sample=True, temperature=0.7, top_p=0.9, top_k=20,
        min_p=0.0, repetition_penalty=1.0, max_new_tokens=256,
    )

print(processor.batch_decode(
    output[:, inputs.input_ids.shape[1]:], skip_special_tokens=True
)[0])

Evaluation and limitations

  • —Packaged-default stress gate: 100/100 responses used first person, with 0 third-person failures and 0 generic refusals.
  • —This model is intended for fictional interaction between adults. It may produce profanity and explicit sexual content.
  • —Generated continuity and boundary handling can fail. Users should review output and restate scene facts when needed.
  • —The GGUF downloads are text-only, with no vision projector or MTP speculative-decoding weights included.
  • —Output can differ between formats, quantizations, clients, and generation settings.

Attribution and release status

Based on Qwen/Qwen3.5-9B. The upstream model is distributed under Apache 2.0; its license is retained in LICENSE-QWEN.

This folder is a local release candidate. Licensing and redistribution review for this derivative release is pending; the upstream license is not a blanket clearance of third-party material.

GGUF runtime: ggml-org/llama.cpp.