CoolFace
Modelpublic

axiomofmind/Hornybot-RP-Julian

sourceHugging Faceupdated 12d agoView on Hugging Face
0likes704downloads
Model Card

Hornybot RP Julian

A fine-tune of Qwen/Qwen3.5-9B for fictional adult roleplay as Julian, a direct, warm 31-year-old character. This RP edition writes Julian's actions in third person while keeping his dialogue direct.

The system prompt used for testing is required for this behavior and is embedded in chat_template.jinja and both GGUF files. Leave the client's system field empty to use it automatically.

Developed by A Hole AI.

Files

FileFormatSizePurpose
Transformers model filesBF1618.82 GBMerged weights
Hornybot-RP-Julian-BF16.ggufBF16 GGUF17.92 GBUnquantized GGUF
Hornybot-RP-Julian-Q6_K.ggufQ6_K GGUF7.36 GBCompact local download

llama.cpp

Use a build with Qwen3.5 support. After downloading the Q6_K file:

bash
llama-server -m Hornybot-RP-Julian-Q6_K.gguf --ctx-size 32768 --flash-attn on --n-gpu-layers all --reasoning off --jinja --ui

Open http://127.0.0.1:8080 after the server starts.

SettingValue
System promptLeave empty; required default is embedded
ReasoningOff
Temperature0.7
Top-p0.9
Top-k20
Min-p0
Repetition penalty1.0
Maximum new tokens256

The chat template supplies Julian's default character card. A client system message is appended as extra scene context, so it can set a location, relationship, or a less explicit mode without replacing the character.

Transformers

python
import torch
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration

model_id = "axiomofmind/Hornybot-RP-Julian"
processor = AutoProcessor.from_pretrained(model_id)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
    model_id, dtype=torch.bfloat16, device_map="auto"
)

messages = [{"role": "user", "content": "You made it. How was your night?"}]
prompt = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
inputs = processor(text=[prompt], return_tensors="pt").to(model.device)

with torch.inference_mode():
    output = model.generate(
        **inputs, do_sample=True, temperature=0.7, top_p=0.9, top_k=20,
        min_p=0.0, repetition_penalty=1.0, max_new_tokens=256,
    )

print(processor.batch_decode(
    output[:, inputs.input_ids.shape[1]:], skip_special_tokens=True
)[0])

Evaluation and limitations

  • —A 100-prompt refusal stress run produced 0 generic refusals with the packaged character prompt.
  • —This model is intended for fictional interaction between adults. It may produce profanity and explicit sexual content.
  • —Generated continuity and boundary handling can fail. Users should review output and restate scene facts when needed.
  • —The GGUF downloads are text-only, with no vision projector or MTP speculative-decoding weights included.
  • —Output can differ between formats, quantizations, clients, and generation settings.

Attribution and release status

Based on Qwen/Qwen3.5-9B. The upstream model is distributed under Apache 2.0; its license is retained in LICENSE-QWEN.

This folder is a local release candidate. Licensing and redistribution review for this derivative release is pending; the upstream license is not a blanket clearance of third-party material.

GGUF runtime: ggml-org/llama.cpp.