fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-50k-instruct
Luna-Protocol-1.5B-Discord-Dialogues-50k-instruct : GGUF
Renamed repo. This model was previously published asLuna-Protocol-1.5B-Discord-Dialogues. The weights are unchanged — the name now spells out the training scale (50k examples) and the base type (Qwen2.5-1.5B-Instruct) so it's identifiable at a glance. If you have the old name saved anywhere (scripts, Modelfiles,llama-cli -hf ...), update it tofox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-50k-instruct.
Luna-Protocol-1.5B-Discord-Dialogues-50k-instruct is a QLoRA fine-tune of Qwen2.5-1.5B-Instruct trained on Discord-Dialogues-Preprocessed-Luna-Protocol (a preprocessed fork of mookiezi/Discord-Dialogues), aimed at reproducing the informal, short-form conversational style of real Discord chat.
Training was done with Unsloth (LoRA, r=16, ~1.18% of parameters trained) on a Kaggle T4, then merged and exported to GGUF.
⚠️ Read the "Recommended usage" section below before judging output quality — with a bare prompt and no priming, this model tends to fall back on Qwen's default assistant tone. A short few-shot prime (shown below) makes a large difference.
This model in context
This card documents the model weights. If what you actually want is a ready-to-run, fully autonomous Discord bot that uses this model as its local LLM — with sleep schedules, typos, hesitation, voice messages, spontaneous messages, anti-spam queueing, and a config.yml that wires few-shot priming in automatically — that's [Luna Protocol](https://github.com/fox3000foxy/luna-protocol-project) on GitHub. The bot downloads and drives this exact GGUF file; everything below (quantizations, priming, limitations) applies directly to it.
Training details
- Base model:
unsloth/Qwen2.5-1.5B-Instruct-bnb-4bit - Method: QLoRA (4-bit),
r=16,lora_alpha=16, target modules:q/k/v/o_proj,gate/up/down_proj - Dataset: ~50,000 examples (subset of the 7.3M-row Discord-Dialogues), filtered to 8–512 tokens, 2–3 epochs
- Trainable params: 18,464,768 / 1,562,179,072 (1.18%)
This is a relatively small-scale fine-tune (50k examples, not the full 7.3M-row dataset) — it shifts the model's tone and register noticeably, but doesn't fully override Qwen's underlying instruction-following behavior. See "Known limitations" below.
Available model files
Recommended usage: few-shot priming
Because the training data (Discord-Dialogues) contains only user/assistant turns and no system-role examples, this model responds only weakly to system prompts alone. What works much better is priming the conversation with a couple of example exchanges in the target style, using the same ChatML structure the model was trained on:
<|im_start|>user
yo whats good<|im_end|>
<|im_start|>assistant
nm just chillin, u<|im_end|>
<|im_start|>user
same tbh, bored af<|im_end|>
<|im_start|>assistant
lol same energy fr<|im_end|>Feed this before the real user turn, then continue the conversation normally. In testing, this consistently produced short, casual, in-character replies (e.g. "suree", "just playing a bit wbu"), versus generic assistant-toned replies (e.g. "Good to know, what's your name?") when using a bare prompt or a verbose instructive system prompt.
A lightweight, non-instructive system prompt (e.g. "you're just chatting with friends on a discord server, nothing formal") can be used in addition to the few-shot prime, but performs poorly on its own without it.
If you'd rather not manage priming by hand, Luna Protocol exposes this as afew_shot_exampleslist inconfig.ymland injects it automatically before every request.
llama.cpp
llama-cli -m Luna-Protocol-1.5B-Fine-Tuned-Qwen2.5.Q8_0.gguf \
--temp 1.0 --top-p 0.9 --top-k 60 --repeat-penalty 1.15 \
-p "<|im_start|>user
yo whats good<|im_end|>
<|im_start|>assistant
nm just chillin, u<|im_end|>
<|im_start|>user
same tbh, bored af<|im_end|>
<|im_start|>assistant
lol same energy fr<|im_end|>
" \
-cnvOr via the HF integration:
llama-cli -hf fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-50k-instruct --jinja<!-- ### Ollama
An Ollama Modelfile is included, using the MESSAGE directive to bake the few-shot prime directly into the model — no manual priming needed at inference time:
FROM Luna-Protocol-1.5B-Fine-Tuned-Qwen2.5.Q8_0.gguf
PARAMETER stop "<|im_end|>"
PARAMETER stop "<|endoftext|>"
PARAMETER temperature 1.0
PARAMETER top_p 0.9
PARAMETER repeat_penalty 1.15
SYSTEM """you're just chatting with friends on a discord server, nothing formal"""
MESSAGE user yo whats good
MESSAGE assistant nm just chillin, u
MESSAGE user same tbh, bored af
MESSAGE assistant lol same energy frollama create luna-protocol -f Modelfile
ollama run luna-protocolKnown limitations
- Weak instruction-following for style directives: asking the model within the system prompt to adopt a specific quirk (e.g. "talk in all lowercase with abbreviations") is not reliably followed — the model tends to keep its own learned tone rather than adapt to fine-grained stylistic instructions.
- Short training run: fine-tuned on ~50k of the 7.3M available rows for 2–3 epochs. A larger-scale run on more of the dataset would likely produce a stronger, more consistent style shift, reducing reliance on few-shot priming.
- Low quantizations degrade style fidelity: Q2K noticeably weakens the learned conversational tone on a model this small; Q4K_M and above preserve it much better.
- Minor context inconsistencies: as expected from a small model, it can contradict earlier turns within a short conversation (e.g. denying playing a game it just discussed).
Credits
- Base model: Qwen2.5-1.5B-Instruct (Qwen team, Alibaba Cloud)
- Training framework: Unsloth
- Dataset: mookiezi/Discord-Dialogues
- Used by: Luna Protocol — the Discord bot this model was trained for
