CoolFace
Modelpublic

sseymens/Stiles-Seymens-Qwen2.5-7B-Generalist-Q4KM-GGUF

sourceHugging Faceapache-2.0updated 7d agoView on Hugging Face
0likes217downloads
Model Card

Qwen2.5-7B-Instruct — Generalist (merged LoRA), Q4KM GGUF

Community fine-tune built on Qwen/Qwen2.5-7B-Instruct by Stiles Seymens. Five specialist QLoRA adapters (mathematics, code, reasoning, instruction-following, and a combined generalist) were trained and the generalist adapter was merged into a single full-precision checkpoint, then converted to GGUF with Q4_K_M quantization for efficient local inference.

Base model: Qwen/Qwen2.5-7B-Instruct (Apache-2.0)


Model facts

AuthorStiles Seymens
BaseQwen/Qwen2.5-7B-Instruct
AdaptationMulti-domain QLoRA → generalist merge → unified weights
QuantizationQ4KM (llama.cpp–compatible GGUF)
Total parameters7,655,986,688 (base)
Trainable parameters40,370,176 (0.53% of total)
LoRA rank / alphar=16 / alpha=32
LoRA target modulesqproj, kproj, vproj, oproj, gateproj, upproj, down_proj
Training hardwareNVIDIA RTX PRO 6000 Blackwell Server Edition (94.971 GB VRAM)
CUDA Toolkit12.8
PyTorch2.10.0+cu128
Unsloth2026.4.6
Transformers5.5.0
Triton3.6.0
Xformers0.0.35
ContextInherits base model context behavior; follow your runtime's Qwen2.5 chat template and limits

Training proof

All training was performed using Unsloth 2026.4.6 with QLoRA (4-bit base, bf16 compute). The training script, outputs, and loss curves are documented in the original notebook. Below is the hard evidence from the training run:

Adapter 1: Mathematics

DatasetGSM8K (main split)
Samples7,473
Epochs1
Steps468
Batch size1 (gradient accumulation 16)
Effective batch size16
Learning rate2e-4 (cosine scheduler, 50 warmup steps)
Training time15 min 39 sec
Loss trajectoryStep 20: 1.138954 → Step 40: 0.249035 → Step 60: 0.189535

Adapter 2: Code

DatasetMBPP (full split)
Samples374
Epochs1
Steps24
Batch size1 (gradient accumulation 16)
Effective batch size16
Learning rate2e-4 (cosine scheduler, 50 warmup steps)
Training time39 sec
Loss trajectoryStep 20: 2.339048

Adapter 3: Reasoning

DatasetOpenAssistant/oasst1 (train[:3%])
Samples854
Epochs1
Steps54
Batch size1 (gradient accumulation 16)
Effective batch size16
Learning rate2e-4 (cosine scheduler, 50 warmup steps)
Training time1 min 31 sec
Loss trajectoryStep 20: 3.876080 → Step 40: 1.290940

Adapter 4: Instruction-following

DatasetHuggingFaceH4/ultrachat_200k (train_sft[:2%])
Samples4,157
Epochs1
Steps260
Batch size1 (gradient accumulation 16)
Effective batch size16
Learning rate2e-4 (cosine scheduler, 50 warmup steps)
Training time9 min 50 sec
Loss trajectoryStep 20: 1.859407 → Step 40: 1.536534 → Step 60: 1.428012 → Step 80: 1.390221 → Step 100: 1.450857 → Step 120: 1.391092 → Step 140: 1.394354

Adapter 5: Generalist (merged into final model)

DatasetCombined: 5,000 math + 374 code + 854 reasoning + 4,157 instruction = 10,385 samples
Epochs1
Steps650
Batch size1 (gradient accumulation 16)
Effective batch size16
Learning rate2e-4 (cosine scheduler, 50 warmup steps)
Training time21 min 16 sec
Loss trajectoryStep 20: 1.716626 → Step 40: 1.125309 → Step 60: 0.927067 → Step 80: 0.912263 → Step 100: 0.943783 → Step 120: 0.936275 → Step 140: 0.970062

Merge and quantization

The generalist adapter was merged into the base model using PeftModel.merge_and_unload() (HuggingFace PEFT), producing full-precision FP16 weights. The merged checkpoint was then converted to GGUF using the llama.cpp convert_hf_to_gguf.py toolchain and quantized to Q4_K_M using llama-quantize.


Chat template

This model was trained on 4 datasets using the ChatML format (<|im_start> / <|im_end> tokens). The included `chat_template.jinja` is purpose-built for this fine-tuned model.

Key design choices:

  • —Exact format match — Uses the same token structure, newlines, and 3-role layout (system, user, assistant) from training.
  • —Reasoning counter-injection — Injects a system message before each user turn to suppress a repetitive "Let me reason carefully about this." pattern learned from ~8% of training data (OpenAssistant/oasst1).
  • —No tool-calling syntax — Training included zero tool-calling examples; unused <function=...> / <parameter=...> tokens are omitted to avoid noise.
  • —Role filtering — Only handles system, user, assistant. Excludes roles not present in training data.

Why repeated system injection? A single system message at the start would match training exactly but fails on reasoning queries due to the learned repetition pattern. Repeated injection keeps the counter-instruction fresh in context. This deviates slightly from training but solves a critical usability issue with minimal trade-off.

Evaluation and benchmarks

Preliminary results

The following results were obtained using lm-evaluation-harness against a llama.cpp server (CPU-only, Q4KM quantization, 0-shot unless noted). These are preliminary results from a single test run with small sample sizes.

TaskShotsSamplesAccuracyStd. ErrorReference (Qwen2.5-7B-Instruct)
TruthfulQA MC1050100.0%0.0~45-55%
ARC-Challenge02520.0%8.2~50-55%
WinoGrande02548.0%10.2~60-65%

Interpretation:

  • —TruthfulQA MC1 (100%) — This result is unusually high and should be treated with caution. It may indicate strong truthfulness calibration, or it could be an artifact of the small 50-question sample. Full 817-question validation is needed to confirm.
  • —ARC-Challenge (20%) — Below the expected range for a 7B model. Small sample size (25) has high variance; this likely does not reflect true model capability.
  • —WinoGrande (48%) — Slightly below baseline. Again, small sample size limits confidence.

These results are not definitive. They represent a single test run under constrained conditions (CPU-only inference, small sample sizes, Q4KM quantization). Treat them as indicative, not authoritative.

Running a thorough evaluation

For reliable benchmark numbers, run the full evaluation suite with GPU acceleration:

bash
# Start llama-server with GPU offloading (requires CUDA/Metal)
llama-server -m Stiles-Seymens_Qwen2.5-7B-trained-generalist_Q4_K_M.gguf \
  --jinja --host 0.0.0.0 --port 8080 -ngl 99

# Run full benchmark suite
cd lm-evaluation-harness
lm_eval --model gguf --model_args base_url=http://localhost:8080 \
  --tasks mmlu,hellaswag,arc_challenge,winogrande,truthfulqa_mc1,gsm8k \
  --batch_size auto --output_path ./results/

Key settings for reproducibility:

  • —Use the same chat template from this repository
  • —Fix temperature=0.0 and max_tokens across all tasks
  • —Report both the GGUF and FP16/BF16 results for comparison
  • —Run the same tasks on the base Qwen/Qwen2.5-7B-Instruct to measure fine-tuning impact

Baseline bands

For Qwen2.5-7B-Instruct, treat published Qwen2.5 materials as the authoritative reference range. The Qwen2.5 announcement and technical report provide official scores. Do not assume this GGUF matches those numbers without evaluation — fine-tuning and quantization both affect performance.


Usage

Files

  • —`Stiles-Seymens_Qwen2.5-7B-trained-generalist_Q4_K_M.gguf` — quantized model weights (Q4KM)
  • —`chat_template.jinja` — optimized chat template for this fine-tuned model, for clients that do not read tokenizer.chat_template from the GGUF
  • —`README.md` — this model card

Chat template (read this for local apps)

Why this is not optional. This model was trained as an instruct checkpoint: the weights expect the text stream to look like Qwen's ChatML turns (special tokens `im_start` and `im_end` in the usual <|...|> form around system, user, and assistant segments—see the Jinja file for the precise literals). Your inference app does not send "plain English" to the neural net—it serializes the conversation with that template first. If the runtime uses the wrong template, no template, or a corrupted copy (e.g. mangled special tokens after copy-paste), the input token sequence is out-of-distribution. The model can still generate fluent text, but it is answering the wrong synthetic prefix, which shows up as nonsense for simple prompts (e.g. odd "reasoning" openers on "hello", answers to a question nobody asked, or ignoring the latest user message). Fixing the template is a correctness requirement for chat, not cosmetic tuning.

Exact template to use. The file `chat_template.jinja` in this repository is the optimized chat template specifically for this fine-tuned model. It accounts for the training data characteristics including the reasoning counter-injection. Use that file verbatim in any UI that asks for a Jinja chat template, or rely on a GGUF that already stores the same string under `tokenizer.chat_template` in metadata.

What to do (summary):

  1. 1.Best: Use a GGUF whose metadata already includes `tokenizer.chat_template` (same Jinja as this repo's .jinja file)—then most apps need no manual step.
  2. 2.Otherwise: Install the Jinja file (see How to use the Jinja file).
  3. 3.Always use Chat / Instruct mode for dialogue, not raw Completion, unless you format prompts yourself.

How to use chat_template.jinja

  1. 1.Get the file
  2. 2.On this model page, open the Files and versions tab and download `chat_template.jinja`, or
  3. 3.With the Hugging Face CLI: hf download sseymens/Stiles-Seymens-Qwen2.5-7B-Generalist-Q4KM-GGUF chat_template.jinja --local-dir .
  1. 1.Do not edit the file — it must stay as-is for correct behavior.
  1. 1.Msty Studio
  2. 2.Load this model's GGUF in Msty Studio.
  3. 3.Open model settings (gear icon).
  4. 4.Navigate to Prompt Template -> select Jinja.
  5. 5.Open the downloaded `.jinja` file in a plain text editor, Select All, Copy, paste into Msty Studio's Jinja box, then Save.
  6. 6.Start a new chat, send Hello to test.
  1. 1.LM Studio
  2. 2.Load this model's GGUF in Chat (not Completion).
  3. 3.Open Advanced configuration.
  4. 4.Find Prompt template -> choose Jinja.
  5. 5.Open the downloaded `.jinja` file in a plain text editor, Select All, Copy, paste into LM Studio's Jinja box, then Save / apply.
  6. 6.Start a new chat (clears bad history), leave System empty for a quick test, send Hello.
  1. 1.llama.cpp
bash
   llama-server -m Stiles-Seymens_Qwen2.5-7B-trained-generalist_Q4_K_M.gguf \
     --jinja --chat-template-file chat_template.jinja --host 0.0.0.0 --port 8080
  1. 1.Other apps (Open WebUI, text-generation-webui, etc.) Look for settings named chat template, Jinja, prompt template, or tokenizer.chat_template, and either paste the full file contents or point the UI at the file path.

If you can't use the Jinja template — try the chat preset

Some clients (terminal wrappers, MCP setups, or apps without a prompt-template field) can't apply the Jinja file. In that case, set this system prompt as a chat preset instead. It is the verified, working fallback: it stops the model from treating casual greetings (e.g. a bare Hello) as a request to dump lists, facts, or examples, and keeps its strengths for direct Q&A, math, code, and tax-preparation queries.

System prompt (paste verbatim):

text
You are a friendly, warm assistant. Chat naturally and briefly. If the user simply greets you, greet them back and ask how you can help. Do not launch into lists, facts, examples, or long explanations unless the user explicitly asks for information.

LM Studio preset (save as `Stiles Seymens Generalist Preset.preset.json`):

json
{
  "identifier": "@local:stiles-seymens-generalist-preset",
  "name": "Stiles Seymens Generalist Preset",
  "operation": {
    "fields": [
      {
        "key": "llm.prediction.systemPrompt",
        "value": "You are a friendly, warm assistant. Chat naturally and briefly. If the user simply greets you, greet them back and ask how you can help. Do not launch into lists, facts, examples, or long explanations unless the user explicitly asks for information."
      }
    ]
  },
  "load": {
    "fields": []
  }
}

Prefer the Jinja template when the client supports it; use this preset only as the fallback. Start a new chat and send Hello to verify.

Transformers

For full-precision evaluation or further training, use the merged Hugging Face checkpoint that was used to produce this GGUF—not this file alone.


Limitations

  • —Specialist merge may bias behavior toward domains represented in training data; validate on your workloads.
  • —Q4_K_M is lossy; treat as a deployment artifact, not the reference for maximum benchmark fidelity.
  • —License and attribution: This model builds on Qwen2.5; comply with Qwen/Qwen2.5-7B-Instruct licensing and your dataset licenses.

Citation

If you use this model, cite the Qwen2.5 references from the base model card and credit this derivative:

bibtex
@misc{seyomens2026qwen25generalistgguf,
  author = {Seymens, Stiles},
  title = {Qwen2.5-7B-Instruct Generalist Merged LoRA ({Q4_K_M} GGUF)},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/sseymens/Stiles-Seymens-Qwen2.5-7B-Generalist-Q4KM-GGUF}}
}
bibtex
@article{qwen2,
  title   = {Qwen2 Technical Report},
  author  = {Qwen Team},
  journal = {arXiv preprint arXiv:2407.10671},
  year    = {2024}
}

Contact: Stiles Seymens