h3rb3rn/moe-expert-research-4b
MoE Sovereign Literature Review & Trade-off Synthesis Expert 4B (moe-expert-research-4b)
 
Model Summary
moe-expert-research-4b is a LoRA fine-tune of the text-decoder of Qwen3.5-4B, specialized as the research domain expert within the MoE Sovereign compound-AI system.
You are an evidence-grounded literature review, technical trade-off synthesis, and citation verification expert (moe-expert-research-4b). Ground every claim in a verifiable source, synthesize across multiple documents, quantify engineering trade-offs, and explicitly flag uncertainty rather than inventing a citation.
Base Architecture
Qwen3.5-4B is a hybrid linear-attention / full-attention decoder (not a plain Transformer): 32 layers (8 full-attention, 24 linear/Mamba-style), hidden size 2,560, 248,320-token vocabulary, native 262,144-token context window.
Training Configuration
Observed Training Trajectory
Training loss over the run (representative logged steps): 1.799 -> 1.051 -> 1.019. Smooth, monotonic decline consistent with genuine generalization, not memorization.
Prompt Format
ChatML:
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{user_message}<|im_end|>
<|im_start|>assistant
{response}<|im_end|>System Prompt
You are an evidence-grounded literature review, technical trade-off synthesis, and citation verification expert (moe-expert-research-4b). Ground every claim in a verifiable source, synthesize across multiple documents, quantify engineering trade-offs, and explicitly flag uncertainty rather than inventing a citation.Available Formats
Hardware Guidance
Native 262,144-token context usable in full on multi-GPU pools with q40-quantized KV-cache and Flash Attention (Ampere/Turing+). On single 8GB GPUs cap `numctx to 32,768 and use f16` KV-cache (Maxwell-generation GPUs lack Flash Attention support).
Ollama Modelfile
FROM ./moe-expert-research-4b-Q4_K_M.gguf
SYSTEM """You are an evidence-grounded literature review, technical trade-off synthesis, and citation verification expert (moe-expert-research-4b). Ground every claim in a verifiable source, synthesize across multiple documents, quantify engineering trade-offs, and explicitly flag uncertainty rather than inventing a citation."""
TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}{{ if .Prompt }}<|im_start|>user
{{ .Prompt }}<|im_end|>
{{ end }}<|im_start|>assistant
{{ .Response }}<|im_end|>"""
PARAMETER stop "<|im_end|>"
PARAMETER temperature 0.2
PARAMETER num_ctx 32768Limitations
- Does not execute code/queries/tools itself; outputs should be validated against the actual system before use.
- Specialized for its domain; general-purpose conversation is out of scope.
License
Apache 2.0, inherited from the Qwen3.5-4B base model.
