CoolFace
Modelpublic

SimplySara/LocoOperator-4B-GGUF

sourceHugging Facemitupdated 7mo agoView on Hugging Face
3likes286downloads
Model Card

This is a static quantization of LocoreMind/LocoOperator-4B, made by SimplySara

ModelSize_GBBPWPPL_QKLD_MeanKLD_MaxTop_P_Match
LocoOperator-4B-BF16.gguf7.49816.019.24309-1.2e-054e-06100.000%
LocoOperator-4B-MXFP4_MOE.gguf3.9868.519.246060.0018352.9823897.518%
LocoOperator-4B-i1-MXFP4_MOE.gguf3.9868.519.246060.0018352.9823897.518%
LocoOperator-4B-Q8_0.gguf3.9868.519.246060.0018352.9823897.518%
LocoOperator-4B-i1-Q8_0.gguf3.9868.519.246060.0018352.9823897.518%
LocoOperator-4B-Q6_K.gguf3.0796.589.279260.006810.568695.526%
LocoOperator-4B-i1-Q6_K.gguf3.0796.589.2950.00607515.994595.857%
LocoOperator-4B-i1-Q5_1.gguf2.8416.079.288590.013642.9883894.135%
LocoOperator-4B-Q5_1.gguf2.8416.079.432220.02267516.345493.161%
LocoOperator-4B-Q5KM.gguf2.6915.759.354570.01702312.394793.635%
LocoOperator-4B-i1-Q5KM.gguf2.6915.759.29650.0131537.7861394.257%
LocoOperator-4B-i1-Q5_0.gguf2.6365.639.422550.01966317.9493.208%
LocoOperator-4B-Q5_0.gguf2.635.629.415210.02340331.401992.839%
LocoOperator-4B-Q5KS.gguf2.635.629.440870.02211913.648392.800%
LocoOperator-4B-i1-Q5KS.gguf2.635.629.287670.0148657.6516993.702%
LocoOperator-4B-Q4_1.gguf2.4185.169.667220.07471815.086187.757%
LocoOperator-4B-i1-Q4_1.gguf2.4185.169.452930.03870713.844490.574%
LocoOperator-4B-Q4KM.gguf2.3264.979.482390.04823615.310590.300%
LocoOperator-4B-i1-Q4KM.gguf2.3264.979.485820.0336813.55191.233%
LocoOperator-4B-IQ4_NL.gguf2.2294.769.608910.05017311.432489.708%
LocoOperator-4B-i1-Q4KS.gguf2.224.749.476030.03984310.055190.557%
LocoOperator-4B-Q4KS.gguf2.224.749.802360.06882115.20988.513%
LocoOperator-4B-i1-IQ4_NL.gguf2.2184.749.502230.0394148.1896490.573%
LocoOperator-4B-i1-Q4_0.gguf2.2134.739.790260.06391512.692888.737%
LocoOperator-4B-Q4_0.gguf2.2074.719.866290.07452713.250187.758%
LocoOperator-4B-IQ4_XS.gguf2.1294.559.621930.05191111.068289.705%
LocoOperator-4B-i1-IQ4_XS.gguf2.1154.529.496870.0400987.0387590.402%
LocoOperator-4B-Q3KL.gguf2.0864.4510.24760.12194427.025784.146%
LocoOperator-4B-i1-Q3KL.gguf2.0864.459.908110.09087415.812286.154%
LocoOperator-4B-Q3KM.gguf1.9334.1310.70210.1578820.204482.662%
LocoOperator-4B-i1-Q3KM.gguf1.9334.139.980570.10270816.824385.354%
LocoOperator-4B-i1-IQ3_M.gguf1.8283.910.16340.13734714.688383.180%
LocoOperator-4B-IQ3_M.gguf1.8283.914.25390.55771319.439767.631%
LocoOperator-4B-IQ3_S.gguf1.7693.7815.06240.61913120.12265.931%
LocoOperator-4B-i1-IQ3_S.gguf1.7693.7810.17550.14206617.002883.139%
LocoOperator-4B-i1-Q3KS.gguf1.7573.7510.88860.17122428.337382.133%
LocoOperator-4B-Q3KS.gguf1.7573.7511.54750.23789530.686879.412%
LocoOperator-4B-i1-IQ3_XS.gguf1.693.6110.36290.16878314.335881.928%
LocoOperator-4B-i1-Q2_K.gguf1.5553.3212.15740.32865218.662275.570%
LocoOperator-4B-i1-IQ3_XXS.gguf1.5553.3211.27950.26344825.25177.569%
LocoOperator-4B-Q2_K.gguf1.5553.3217.1530.71359616.394664.880%
LocoOperator-4B-i1-Q2KS.gguf1.4563.1113.17090.45012518.382671.231%
LocoOperator-4B-i1-IQ2_M.gguf1.4093.0114.08570.54476418.561867.933%
LocoOperator-4B-i1-IQ2_S.gguf1.322.8215.07170.62118924.098165.722%
LocoOperator-4B-i1-IQ2_XS.gguf1.2612.6916.82770.75033619.212863.162%
LocoOperator-4B-i1-IQ2_XXS.gguf1.1612.4827.59881.3214414.680752.522%
LocoOperator-4B-i1-IQ1_M.gguf1.052.2449.09781.932316.594744.067%
LocoOperator-4B-i1-IQ1_S.gguf0.9832.1139.9513.0327416.094728.387%

<div align="center"> <img src="assets/loco_operator.png" width="55%" alt="LocoOperator" /> </div>

<br>

<div align="center">

![MODEL](https://huggingface.co/LocoreMind/LocoOperator-4B) ![GGUF](https://huggingface.co/LocoreMind/LocoOperator-4B-GGUF) ![Blog](https://locoremind.com/blog/loco-operator) ![GitHub](https://github.com/LocoreMind/LocoOperator) ![Colab](https://colab.research.google.com/github/LocoreMind/LocoOperator/blob/main/LocoOperator_4B.ipynb)

</div>

Introduction

LocoOperator-4B is a 4B-parameter tool-calling agent model trained via knowledge distillation from Qwen3-Coder-Next inference traces. It specializes in multi-turn codebase exploration — reading files, searching code, and navigating project structures within a Claude Code-style agent loop. Designed as a local sub agent, it runs via llama.cpp at zero API cost.

LocoOperator-4B
Base ModelQwen3-4B-Instruct-2507
Teacher ModelQwen3-Coder-Next
Training MethodFull-parameter SFT (distillation)
Training Data170,356 multi-turn conversation samples
Max Sequence Length16,384 tokens
Training Hardware4x NVIDIA H200 141GB SXM5
Training Time~25 hours
FrameworkMS-SWIFT

Key Features

  • —Tool-Calling Agent: Generates structured <tool_call> JSON for Read, Grep, Glob, Bash, Write, Edit, and Task (subagent delegation)
  • —100% JSON Validity: Every tool call is valid JSON with all required arguments — outperforming the teacher model (87.6%)
  • —Local Deployment: GGUF quantized, runs on Mac Studio via llama.cpp at zero API cost
  • —Lightweight Explorer: 4B parameters, optimized for fast codebase search and navigation
  • —Multi-Turn: Handles conversation depths of 3–33 messages with consistent tool-calling behavior

Performance

Evaluated on 65 multi-turn conversation samples from diverse open-source projects (scipy, fastapi, arrow, attrs, gevent, gunicorn, etc.), with labels generated by Qwen3-Coder-Next.

Core Metrics

MetricScore
Tool Call Presence Alignment100% (65/65)
First Tool Type Match65.6% (40/61)
JSON Validity100% (76/76)
Argument Syntax Correctness100% (76/76)

The model perfectly learned when to use tools vs. when to respond with text (100% presence alignment). Tool type mismatches are between semantically similar tools (e.g. Grep vs Read) — different but often valid strategies.

Tool Distribution Comparison

<div align="center"> <img src="assets/tool_distribution.png" width="80%" alt="Tool Distribution Comparison" /> </div>

JSON & Argument Syntax Correctness

ModelJSON ValidArgument Syntax Valid
LocoOperator-4B76/76 (100%)76/76 (100%)
Qwen3-Coder-Next (teacher)89/89 (100%)78/89 (87.6%)
LocoOperator-4B achieves perfect structured output. The teacher model has 11 tool calls with missing required arguments (empty arguments: {}).

Quick Start

python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "LocoreMind/LocoOperator-4B"

# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto"
)

# prepare the messages
messages = [
    {
        "role": "system",
        "content": "You are a read-only codebase search specialist.\n\nCRITICAL CONSTRAINTS:\n1. STRICTLY READ-ONLY: You cannot create, edit, delete, move files, or run any state-changing commands. Use tools/bash ONLY for reading (e.g., ls, find, cat, grep).\n2. EFFICIENCY: Spawn multiple parallel tool calls for faster searching.\n3. OUTPUT RULES: \n   - ALWAYS use absolute file paths.\n   - STRICTLY NO EMOJIS in your response.\n   - Output your final report directly. Do not use colons before tool calls.\n\nENV: Working directory is /Users/developer/workspace/code-analyzer (macOS, zsh)."
    },
    {
        "role": "user",
        "content": "Analyze the Black codebase at `/Users/developer/workspace/code-analyzer/projects/black`.\nFind and explain:\n1. How Black discovers config files.\n2. The exact search order for config files.\n3. Supported config file formats.\n4. Where this configuration discovery logic lives in the codebase.\n\nReturn a comprehensive answer with relevant code snippets and absolute file paths."
    }
]

# prepare the model input
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

# conduct text completion
generated_ids = model.generate(
    **model_inputs,
    max_new_tokens=512,
)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist()

content = tokenizer.decode(output_ids, skip_special_tokens=True)
print(content)

Local Deployment

For GGUF quantized deployment with llama.cpp, hybrid proxy routing, and batch analysis pipelines, refer to our GitHub repository.

Training Details

ParameterValue
Base modelQwen3-4B-Instruct-2507
Teacher modelQwen3-Coder-Next
MethodFull-parameter SFT
Training data170,356 samples
Hardware4x NVIDIA H200 141GB SXM5
ParallelismDDP (no DeepSpeed)
PrecisionBF16
Epochs1
Batch size2/GPU, gradient accumulation 4 (effective batch 32)
Learning rate2e-5, warmup ratio 0.03
Max sequence length16,384 tokens
Templateqwen3_nothinking
FrameworkMS-SWIFT
Training time~25 hours
CheckpointStep 2524

Known Limitations

  • —First-tool-type match is 65.6% — the model sometimes picks a different (but not necessarily wrong) tool than the teacher
  • —Tends to under-generate parallel tool calls compared to the teacher (76 vs 89 total calls across 65 samples)
  • —Preference for Bash over Read may indicate the model defaults to shell commands where file reads would be more appropriate
  • —Evaluated on 65 samples only; larger-scale evaluation needed

License

MIT

Acknowledgments

  • —Qwen Team for the Qwen3-4B-Instruct-2507 base model
  • —MS-SWIFT for the training framework
  • —llama.cpp for efficient local inference
  • —Anthropic for the Claude Code agent loop design that inspired this work