CoolFace
Modelpublic

h3rb3rn/moe-expert-precision-4b

sourceHugging Faceapache-2.0updated 12d agoView on Hugging Face
3likes140downloads
Model Card

MoE Sovereign Formal-Reasoning & Numerical-Precision Expert 4B (moe-expert-precision-4b)

![License: Apache 2.0](https://opensource.org/licenses/Apache-2.0) ![Base Model: Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B)


Model Summary

moe-expert-precision-4b is a LoRA fine-tune of the text-decoder of Qwen3.5-4B, specialized as the precision domain expert within the MoE Sovereign compound-AI system.

You are a formal-reasoning, SMT constraint formulation, and numerical-precision expert (moe-expert-precision-4b). Decompose quantitative problems into explicit, tool-verifiable calculation steps and always ground exact numbers in a deterministic calculation tool rather than estimating in prose. Never guess a numeric result you can compute exactly; show your work and validate the output before presenting it.

Base Architecture

Qwen3.5-4B is a hybrid linear-attention / full-attention decoder (not a plain Transformer): 32 layers (8 full-attention, 24 linear/Mamba-style), hidden size 2,560, 248,320-token vocabulary, native 262,144-token context window.

Training Configuration

ParameterValue
MethodLoRA (rank 16, alpha 32, dropout 0.05), targeting q/k/v/o_proj + gate/up/down_proj
Trainable parameters21,233,664 of 4,226,984,960 (0.50%)
Epochs3
Effective batch size128 (micro-batch 4 x 8 GPUs x grad-accum 4)
Learning rate1.5e-5
Training sequence length4,096 tokens
Optimizer shardingDeepSpeed ZeRO-2, bf16
ComputeEuroHPC LUMI-G, 8x AMD Instinct MI250X GCDs, ROCm
Training examples5,879 curated instruction/response pairs

Observed Training Trajectory

Training loss over the run (representative logged steps): 0.9714 -> 0.5867 -> 0.5562. Smooth, monotonic decline consistent with genuine generalization, not memorization.

Prompt Format

ChatML:

<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{user_message}<|im_end|>
<|im_start|>assistant
{response}<|im_end|>

System Prompt

You are a formal-reasoning, SMT constraint formulation, and numerical-precision expert (moe-expert-precision-4b). Decompose quantitative problems into explicit, tool-verifiable calculation steps and always ground exact numbers in a deterministic calculation tool rather than estimating in prose. Never guess a numeric result you can compute exactly; show your work and validate the output before presenting it.

Available Formats

FileNotes
moe-expert-precision-4b-Q4_K_M.ggufRecommended for single/multi-GPU deployment
moe-expert-precision-4b-Q8_0.ggufHigher-fidelity reference quantization

Hardware Guidance

Native 262,144-token context usable in full on multi-GPU pools with q40-quantized KV-cache and Flash Attention (Ampere/Turing+). On single 8GB GPUs cap `numctx to 32,768 and use f16` KV-cache (Maxwell-generation GPUs lack Flash Attention support).

Ollama Modelfile

dockerfile
FROM ./moe-expert-precision-4b-Q4_K_M.gguf
SYSTEM """You are a formal-reasoning, SMT constraint formulation, and numerical-precision expert (moe-expert-precision-4b). Decompose quantitative problems into explicit, tool-verifiable calculation steps and always ground exact numbers in a deterministic calculation tool rather than estimating in prose. Never guess a numeric result you can compute exactly; show your work and validate the output before presenting it."""
TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}{{ if .Prompt }}<|im_start|>user
{{ .Prompt }}<|im_end|>
{{ end }}<|im_start|>assistant
{{ .Response }}<|im_end|>"""
PARAMETER stop "<|im_end|>"
PARAMETER temperature 0.2
PARAMETER num_ctx 32768

Limitations

  • Does not execute code/queries/tools itself; outputs should be validated against the actual system before use.
  • Specialized for its domain; general-purpose conversation is out of scope.

License

Apache 2.0, inherited from the Qwen3.5-4B base model.