CoolFace
Modelpublic

Rewnozom/Rewnozom-GGUF

sourceHugging Faceapache-2.0updated 20d agoView on Hugging Face
1likes25kdownloads
Model Card

Rewnozom-GGUF

GGUF conversion of Rewnozom/Rewnozom, derived from Qwen/Qwen2.5-7B-Instruct-1M.

  • Original model: Rewnozom/Rewnozom
  • Converted model: Rewnozom/Rewnozom-GGUF
  • Format: GGUF
  • Runtime targets: llama.cpp, Ollama, LM Studio, Jan, and other GGUF loaders

Personal recommendation: start with the two hybrid MX profiles, MX-k_quants and MX-legacy-quants, because they keep the sensitive tensors higher while leaving the rest on a compact base quant.

Available Quantizations

Only quantizations with an actual .gguf file in this workspace are listed.

QuantFileApprox size
F32F32/Rewnozom-1M-LR.F32.gguf28.38 GB
F16F16/Rewnozom-1M-LR.F16.gguf14.19 GB
BF16BF16/Rewnozom-1M-LR.BF16.gguf14.19 GB
MX-k_quantsMX-k_quants/Rewnozom-1M-LR.Q4_K_M.k_quants_hybrid_q4km_q5_q6.gguf4.62 GB
MX-legacy-quantsMX-legacy-quants/Rewnozom-1M-LR.Q4_0.legacy_hybrid_q40_q50_q80.gguf5.05 GB
MXFP4_MOEextra_current_llama_quants/MXFP4_MOE/Rewnozom-1M-LR.MXFP4_MOE.gguf7.54 GB
Q1_0extra_current_llama_quants/Q1_0/Rewnozom-1M-LR.Q1_0.gguf1.35 GB
Q2_0extra_current_llama_quants/Q2_0/Rewnozom-1M-LR.Q2_0.gguf2.42 GB
Q2_Kk_quants/Q2_K/Rewnozom-1M-LR.Q2_K.gguf2.81 GB
Q3_K_Lk_quants/Q3_K_L/Rewnozom-1M-LR.Q3_K_L.gguf3.81 GB
Q3_K_Mk_quants/Q3_K_M/Rewnozom-1M-LR.Q3_K_M.gguf3.55 GB
Q3_K_Sk_quants/Q3_K_S/Rewnozom-1M-LR.Q3_K_S.gguf3.25 GB
Q4_K_Mk_quants/Q4_K_M/Rewnozom-1M-LR.Q4_K_M.gguf4.36 GB
Q4_K_Sk_quants/Q4_K_S/Rewnozom-1M-LR.Q4_K_S.gguf4.15 GB
Q5_K_Mk_quants/Q5_K_M/Rewnozom-1M-LR.Q5_K_M.gguf5.07 GB
Q5_K_Sk_quants/Q5_K_S/Rewnozom-1M-LR.Q5_K_S.gguf4.95 GB
Q6_Kk_quants/Q6_K/Rewnozom-1M-LR.Q6_K.gguf5.82 GB
Q4_0legacy_quants/Q4_0/Rewnozom-1M-LR.Q4_0.gguf4.13 GB
Q4_1legacy_quants/Q4_1/Rewnozom-1M-LR.Q4_1.gguf4.54 GB
Q5_0legacy_quants/Q5_0/Rewnozom-1M-LR.Q5_0.gguf4.95 GB
Q5_1legacy_quants/Q5_1/Rewnozom-1M-LR.Q5_1.gguf5.36 GB
Q8_0legacy_quants/Q8_0/Rewnozom-1M-LR.Q8_0.gguf7.54 GB
TQ1_0t_quants/TQ1_0/Rewnozom-1M-LR.TQ1_0.gguf1.99 GB
TQ2_0t_quants/TQ2_0/Rewnozom-1M-LR.TQ2_0.gguf2.28 GB

Quantization Groups

GroupPurpose
F32Full float32 GGUF reference conversion. Highest precision, largest file.
F16Float16 GGUF reference conversion. Good baseline for further quantization.
BF16BFloat16 GGUF reference conversion for runtimes that prefer BF16.
extra_current_llama_quantsAdditional current llama.cpp-compatible quant types.
k_quantsK-quant family, usually the best default family for local use.
legacy_quantsOlder GGUF quant family for compatibility and comparison.
t_quantsTernary/low-bit quant family for very small local deployments.
ollama_modelfilesGenerated Ollama Modelfiles using the base system prompt from sp.md.

Recommended Starting Points

Use caseQuant
Personal recommendation, K-family hybridMX-k_quants
Personal recommendation, legacy hybridMX-legacy-quants
Best quality among compact K-quantsQ6_K
Balanced defaultQ4_K_M
Smaller memory footprintQ3_K_M or Q3_K_S
Very small local testQ2_K, Q2_0, TQ2_0, or Q1_0
Legacy compatibility checkQ4_0, Q4_1, Q5_0, Q5_1, Q8_0

Actual quality and speed depend on runtime, CPU/GPU offload, context length, and prompt workload. Validate the target quant against your real tasks before using it as a default.

llama.cpp

Run directly from a local GGUF file:

bash
llama-cli -m k_quants/Q4_K_M/Rewnozom-1M-LR.Q4_K_M.gguf \
  -p "Review this implementation plan for missing constraints."

Start an OpenAI-compatible local server:

bash
llama-server -m k_quants/Q4_K_M/Rewnozom-1M-LR.Q4_K_M.gguf

Dataset Context

The model is associated with a synthetic reasoning/control-plane dataset family covering:

  • boolean CSP logic
  • branch-dependent task DAGs
  • ordering and plan repair
  • multi-hop forward inference
  • request decomposition
  • context relevance
  • deterministic state transitions
  • memory lifecycle
  • retrieval/navigation policy
  • executor routing
  • tool execution
  • permissions
  • multi-agent orchestration
  • result validation
  • retry/escalation
  • composite execution kernel behavior

The dataset design uses deterministic formal worlds, double oracle checks, structural dedupe, and machine-verifiable answers rather than synthetic prose chain-of-thought.

Limitations

  • GGUF quantization changes numerical behavior compared with the source model.
  • Lower-bit quants trade quality for memory and speed.
  • Permission enforcement, destructive actions, and state mutation should remain controlled by deterministic application logic.

Attribution

This conversion is based on Rewnozom/Rewnozom, which is derived from Qwen/Qwen2.5-7B-Instruct-1M and follows the Apache 2.0 license.

page:

Base model: https://huggingface.co/Qwen/Qwen2.5-7B-Instruct-1M

Rewnozom/Rewnozom: https://huggingface.co/Rewnozom/Rewnozom

Rewnozom/Rewnozom-GGUF: https://huggingface.co/Rewnozom/Rewnozom-GGUF

Ollama: https://ollama.com/tobraa92/Rewnozom

Portfolio: https://tobiasraanaes.se/