koshuro/Qwen-DeepSeek-1.5B-Agentic-distill
0365
Qwen-DeepSeek-1.5B-Agentic-distill : GGUF
This model was finetuned on 1 billion tokens of filtered reasoning traces, mainly consisting off tool use, agentic work, and repository work.
Benchmark results VS a similar sized model.

Example usage:
- For text only LLMs:
llama-cli -hf koshuro/Qwen-DeepSeek-1.5B-Agentic-distill --jinja - For multimodal models:
llama-mtmd-cli -hf koshuro/Qwen-DeepSeek-1.5B-Agentic-distill --jinja
Available Model files:
deepseek-r1-distill-qwen-1.5b.Q3_K_M.ggufdeepseek-r1-distill-qwen-1.5b.F16.ggufdeepseek-r1-distill-qwen-1.5b.Q2_K_L.ggufdeepseek-r1-distill-qwen-1.5b.Q4_K_M.ggufdeepseek-r1-distill-qwen-1.5b.Q6_K.ggufdeepseek-r1-distill-qwen-1.5b.Q8_0.gguf
Note
The model's BOS token behavior was adjusted for GGUF compatibility. This was trained 2x faster with Unsloth <img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/>
