CoolFace
Modelpublic

notshekhar/markdown-1

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes53downloads
Model Card

markdown-1

VibeThinker-3B fine-tuned (LoRA, merged) for tool calling + long agent traces.

This repo contains the merged fp16 weights plus ready-to-run GGUF quants for llama.cpp / Ollama / LM Studio.

FileSizeUse
markdown-1-Q4_K_M.gguf~1.9 GBsmaller / faster, great default
markdown-1-Q8_0.gguf~3.3 GBhigher fidelity
model-*.safetensors~6.2 GBmerged fp16 (vLLM / transformers)

LoRA adapter only: `notshekhar/vibethinker-finetuned-tool`.

Run with llama.cpp

bash
llama-cli -hf notshekhar/markdown-1:Q4_K_M -p "Hello"
# or local:
llama-cli -m markdown-1-Q4_K_M.gguf -p "Hello"

Run with Ollama

bash
# Modelfile
printf 'FROM ./markdown-1-Q4_K_M.gguf\n' > Modelfile
ollama create markdown-1 -f Modelfile
ollama run markdown-1

Base reasoning model uses <think> traces and ChatML (<|im_start|>) with tool-calling via <tool_call> / <tool_response> blocks (see chat_template.jinja).