CoolFace
Modelpublic

KoarAI/LFM2.5-350M-Thinking

sourceHugging Faceapache-2.0updated 29d agoView on Hugging Face
0likes707downloads
Model Card

<div align="center">

<img src="https://huggingface.co/KoarAI/LFM2.5-350M-Thinking/resolve/main/banner.png" alt="KoarAI LFM2.5-350M Thinking Banner" width="100%" style="border-radius: 12px; box-shadow: 0 4px 20px rgba(0,0,0,0.3);"/>

๐Ÿจ KoarAI / LFM2.5-350M-Thinking

![License: Apache 2.0](https://opensource.org/licenses/Apache-2.0) ![Model Revision](https://huggingface.co/KoarAI/LFM2.5-350M-Thinking) ![Fine-Tuning: 100% Full Weights](https://huggingface.co/KoarAI/LFM2.5-350M-Thinking) ![Parameters](https://huggingface.co/KoarAI/LFM2.5-350M-Thinking) ![GGUF Quantized](https://huggingface.co/KoarAI/LFM2.5-350M-Thinking-GGUF)

</div>

๐Ÿ“Œ Release Note: Model Code 0002 (Weight Architecture Update)

[!IMPORTANT] Model Code: 0002 This version underwent a comprehensive 100% Full Parameter Fine-Tuning across 9 epochs with a cosine learning rate scheduler. It integrates an expanded multi-teacher dataset (Reasoning CoT + DeepSeek-V4-Pro Agentic + MMLU-Pro + AIME 2026 Mathematics) and strict syntactic normalization for <think> ... </think> blocks. ๐Ÿš€ KoarAI Release & Versioning Policy: Starting from the upcoming release (0003 and beyond), rather than overwriting existing models, each new iteration will be released into its own dedicated repository (e.g., KoarAI/LFM2.5-350M-Thinking-v3, KoarAI/LFM2.5-350M-Thinking-RU, etc.).

๐ŸŒŸ Overview

`KoarAI/LFM2.5-350M-Thinking` (Code: 0002) is an ultra-compact, high-efficiency language model featuring native Chain-of-Thought (CoT) reasoning capabilities.

Built upon the state-of-the-art Liquid Foundation Model architecture (LiquidAI/LFM2.5-350M), this model was trained using 100% Full Parameter Fine-Tuning on a balanced blend of distilled reasoning traces from frontier models:

  • โ€”`Qwen 3.8 Max`
  • โ€”`GLM 5.2`
  • โ€”`Kimi K3`
  • โ€”`DeepSeek-V4-Pro 0813 Agentic`
  • โ€”`MMLU-Pro & AIME 2026 Mathematics`

Despite having only 350 Million parameters, the model demonstrates strong multi-step logic, mathematical deduction, and structured problem-solving inside native <think> ... </think> blocks.


๐Ÿ’ก Native Thinking Mode

The model natively reasons before outputting its final response:

text
<|im_start|>user
Solve: 32 + 32 - 42<|im_end|>
<|im_start|>assistant
<think>
1. Evaluate 32 + 32 = 64.
2. Subtract 42 from 64: 64 - 42 = 22.
</think>
\boxed{22}<|im_end|>

โšก Quickstart

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "KoarAI/LFM2.5-350M-Thinking"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto",
    trust_remote_code=True
)

messages = [
    {"role": "user", "content": "How many 'r' in strawberry?"}
]

prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

outputs = model.generate(
    **inputs,
    max_new_tokens=512,
    temperature=0.6,
    top_p=0.9,
    do_sample=True,
    pad_token_id=tokenizer.eos_token_id
)

print(tokenizer.decode(outputs[0][inputs['input_ids'].shape[1]:], skip_special_tokens=False))

๐Ÿ“ฆ GGUF & Quantization

Official quantized GGUF versions (FP16, Q8_0, Q5_K_M, Q4_K_M, Q4_0) for llama.cpp, Ollama, and LM Studio are available at: ๐Ÿ‘‰ **`KoarAI/LFM2.5-350M-Thinking-GGUF`**


๐Ÿจ Maintained by KoarAI Lab

Released for the open-source AI community by KoarAI.