CoolFace
Modelpublic

AX1Y2JP/LFM2.5-VL-3B-heretic-GGUF

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
2likes804downloads
Model Card

base model:https://huggingface.co/heretic-org/LFM2.5-VL-3B-heretic

This is a decensored version of LiquidAI/LFM2.5-VL-3B, made using Heretic v1.4.0

[!TIP] This model is reproducible! See the README in the reproduce directory for more information.

Abliteration parameters

ParameterValue
direction_indexper layer
attn.o_proj.max_weight1.20
attn.o_proj.max_weight_position22.59
attn.o_proj.min_weight0.60
attn.o_proj.min_weight_distance11.46
mlp.down_proj.max_weight1.39
mlp.down_proj.max_weight_position23.51
mlp.down_proj.min_weight0.94
mlp.down_proj.min_weight_distance13.07

Performance

MetricThis modelOriginal model ([LiquidAI/LFM2.5-VL-3B](https://huggingface.co/LiquidAI/LFM2.5-VL-3B))
KL divergence0.05530 (by definition)
Refusals6/100100/100

<div align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png" alt="Liquid AI" style="width: 100%; max-width: 100%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;" /> <div style="display: flex; justify-content: center; gap: 0.5em; margin-bottom: 1em;"> <a href="https://playground.liquid.ai/"><strong>Try LFM</strong></a> • <a href="https://docs.liquid.ai/lfm/getting-started/welcome"><strong>Docs</strong></a> • <a href="https://leap.liquid.ai/"><strong>LEAP</strong></a> • <a href="https://discord.com/invite/liquid-ai"><strong>Discord</strong></a> </div> </div>

LFM2.5-VL-3B

LFM2.5-VL-3B is a multimodal variant of LFM2.5, a family of hybrid models designed for on-device deployment. It builds on LFM2-VL-3B with further mid- and post-training. LFM2.5-VL-3B can process both text and images, and uses the LFM2.5-2.6B language model as its backbone, combined with a SigLIP2 NaFlex vision encoder.

  • —Better grounding: Improved grounding and object detection with natural language queries.
  • —Better OCR: Full page OCR with layout annotation. See layout annotation format for more information.
  • —Efficient inference: 228 tok/s on an Apple M5 Max and 116 tok/s on an AMD Ryzen AI Max+ 395, in under 3.3 GB of memory.

Find more information about LFM2.5-VL-3B in our release post.

lfm2_5_vl_3b_task_group_averages

[!NOTE] 💻 Demos: Try LFM2.5-VL-3B's vision understanding capabilities in a Hugging Face space without any setup: [Vision-capable chat in your browser](https://huggingface.co/spaces/LiquidAI/LFM2.5-VL-3B-WebGPU): allows you to upload images or use the webcam to capture stills and let the model interact with them.

Model Details

ModelDescription
[LFM2.5&#8209;VL&#8209;3B](https://huggingface.co/LiquidAI/LFM2.5-VL-3B)Original checkpoint in native format. Best for fine-tuning and inference with HF Transformers, vLLM and SGLang
[LFM2.5&#8209;VL&#8209;3B&#8209;GGUF](https://huggingface.co/LiquidAI/LFM2.5-VL-3B-GGUF)Quantized GGUF exports of the original checkpoint. Best for CPU inference with reduced memory usage with llama.cpp
[LFM2.5&#8209;VL&#8209;3B&#8209;ONNX](https://huggingface.co/LiquidAI/LFM2.5-VL-3B-ONNX)Quantized ONNX exports for cross-platform deployment. Enables hardware-accelerated inference across diverse environments (cloud, edge, mobile). See the demo.
[LFM2.5-VL-3B-MLX](https://huggingface.co/LiquidAI/LFM2.5-VL-3B-MLX-8bit)Quantized MLX exports for Apple Silicon. Optimized for fast inference on Mac devices using the mlx-vlm framework.
  • —LM Backbone: LFM2.5-2.6B
  • —Vision encoder: SigLIP2 NaFlex shape‑optimized 400M
  • —Vocabulary size: 128,000
  • —Context length: 32,768 tokens
  • —Languages: English, Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Spanish, Vietnamese, Thai, Indonesian, Hindi, Russian, Polish
  • —Native resolution processing: Uses SigLIP2's NaFlex; large images are split into non-overlapping 512×512 patches and a resized whole-image thumbnail.
  • —Generation parameters:
  • —text: temperature=0.2, top_k=50, repetition_penalty=1.0
  • —vision: Use the processor_config.json file.

We recommend using it for single-turn, high-throughput, low-latency tasks; for example, for near-realtime object detection in automotive applications, batch processing scanned documents with OCR with layout information for turning PDFs into searchable text, or for on-device translation of menus and road signs into your native language.

It is not recommended for long-context, reasoning-intensive tasks, such as visual web design, or answering highly technical questions about blueprints.

Chat Template

LFM2.5 uses a ChatML-like format. See the Chat Template documentation for details. Example:

<|startoftext|><|im_start|>system
You are a helpful assistant trained by Liquid AI.<|im_end|>
<|im_start|>user
What species is in this picture?<image><|im_end|>
<|im_start|>assistant

You can use `tokenizer.apply_chat_template()` to format your messages automatically.

[!TIP] Note: The apply_chat_template() method automatically inserts the <image> tag for each image in your message. Do not include <image> in your message content.

Inference

LFM2.5-VL is supported by many inference frameworks. See the Inference documentation for the full list.

NameDescriptionDocsNotebook
TransformersSimple inference with direct access to model internals.<a href="https://docs.liquid.ai/lfm/inference/transformers#vision-models">Link</a><a href="https://colab.research.google.com/drive/1WVQpf4XrHgHFkP0FnlZfx2nK8PugvQNZ?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHabLXysEu2E.png" width="110" alt="Colab link"></a>
vLLMHigh-throughput production deployments with GPU.<a href="https://docs.liquid.ai/deployment/gpu-inference/vllm#vision-models">Link</a><a href="https://colab.research.google.com/drive/1sUfQlqAvuAVB4bZ6akYVQPGmHtTDUNpF?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHabLXysEu2E.png" width="110" alt="Colab link"></a>
SGLangHigh-throughput production deployments with GPU.<a href="https://docs.liquid.ai/deployment/gpu-inference/sglang#vision-models">Link</a><a href="https://colab.research.google.com/drive/1qJlAFag223yFOZGzuMIkYUFhybM9ao5g?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHabLXysEu2E.png" width="110" alt="Colab link"></a>
llama.cppCross-platform inference with CPU offloading.<a href="https://docs.liquid.ai/lfm/inference/llama-cpp#vision-models">Link</a><a href="https://colab.research.google.com/drive/1q2PjE6OAahakRlkTNJGYL32MsdUcj7b?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHab_LXysEu2E.png" width="110" alt="Colab link"></a>

Quick start

Quick start with Transformers (compatible with transformers>=5.0.0):

You will need torch, transformers, and torchvision.

python
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor

model_id = "LiquidAI/LFM2.5-VL-3B"

model = AutoModelForImageTextToText.from_pretrained(model_id, dtype=torch.bfloat16)
processor = AutoProcessor.from_pretrained(model_id)

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://placecats.com/300/200"},
            {"type": "text", "text": "Describe this image."},
        ],
    }
]

inputs = processor.apply_chat_template(
    messages,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)

with torch.inference_mode():
    output_ids = model.generate(
        **inputs,
        do_sample=True,
        temperature=0.2,
        top_k=50,
        repetition_penalty=1.0,
        max_new_tokens=256,
    )

generated_ids = output_ids[:, inputs["input_ids"].shape[1] :]
print(processor.batch_decode(generated_ids, skip_special_tokens=True)[0])

Tool Use

LFM2.5-VL-3B supports function calling in four steps:

  1. 1.Function definition: Provide the list of tools as a JSON object in the system prompt, or use `tokenizer.apply_chat_template()` with tools=....
  2. 2.Function call: By default, LFM2.5 writes Pythonic function calls (a Python list between <|tool_call_start|> and <|tool_call_end|> special tokens), as the assistant answer.
  3. 3.Function execution: Execute the call and return the result with the tool role.
  4. 4.Final answer: LFM2.5 interprets the tool output and returns a plain-text answer addressing the original prompt.

See the Tool Use documentation for the full guide. Example:

<|startoftext|><|im_start|>system
List of tools: [{"name": "get_candidate_status", "description": "Retrieves the current status of a candidate in the recruitment process", "parameters": {"type": "object", "properties": {"candidate_id": {"type": "string", "description": "Unique identifier for the candidate"}}, "required": ["candidate_id"]}}]<|im_end|>
<|im_start|>user
What is the current status of candidate ID 12345?<|im_end|>
<|im_start|>assistant
<|tool_call_start|>[get_candidate_status(candidate_id="12345")]<|tool_call_end|>Checking the current status of candidate ID 12345.<|im_end|>
<|im_start|>tool
[{"candidate_id": "12345", "status": "Interview Scheduled", "position": "Clinical Research Associate", "date": "2023-11-20"}]<|im_end|>
<|im_start|>assistant
The candidate with ID 12345 is currently in the "Interview Scheduled" stage for the position of Clinical Research Associate, with an interview date set for 2023-11-20.<|im_end|>

Layout Annotation Format

LFM2.5-VL-3B can do OCR with layout annotation. The layout annotation is a list of regions, each with a label, bounding box, and content. The format is:

text
image_index=<n> <label> [xmin, ymin, xmax, ymax]
<content>

image_index=<n> <label> [xmin, ymin, xmax, ymax]
<content>

image_index=<n> <label> [xmin, ymin, xmax, ymax]
<content>

...

where:

  • —image_index is the zero-based index of image
  • —<label> is one of these layout labels:
  • —text
  • —title
  • —list
  • —table
  • —table_caption
  • —table_footnote
  • —image
  • —image_block
  • —image_caption
  • —image_footnote
  • —chart
  • —equation
  • —formula_number
  • —code
  • —code_caption
  • —algorithm
  • —aside_text
  • —ref_text
  • —phonetic
  • —page_header
  • —page_footer
  • —page_number
  • —page_footnote
  • —[xmin, ymin, xmax, ymax] are normalized integer coordinates in [0, 1000], same as our grounding format.
  • —<content> is the region's content:
  • —plain text for text regions
  • —LaTeX for equations
  • —OTSL (Optimized Table Structure Language, introduced here by IBM) for tables
  • —a short description for images and charts

There will be a blank line between each region.

To prompt the model to generate this structured output, use a system or user prompt that includes this:

text
Parse this document into its layout regions. The pages are provided as images in reading order. For every region, in reading order across all pages, output a header line immediately followed by the region's content:

image_index=<n> <label> [xmin, ymin, xmax, ymax]
<content>

where:
- image_index is the zero-based index of the page image the region appears on (0 for the first image, 1 for the second, and so on)
- <label> is one of these layout labels: text, title, list, table, table_caption, table_footnote, image, image_block, image_caption, image_footnote, chart, equation, formula_number, code, code_caption, algorithm, aside_text, ref_text, phonetic, page_header, page_footer, page_number, page_footnote
- [xmin, ymin, xmax, ymax] are normalized integer coordinates in [0, 1000]
- <content> is the region's content: plain text for text regions, LaTeX for equations, OTSL for tables, and a short description for images and charts

Separate each region block with one blank line. Return only the parsed regions.

Note that the layout annotation format is still experimental: it may change, may be unreliable, and may not be trivial to parse. We encourage users to try it out and provide feedback!

Fine-Tuning

We recommend fine-tuning LFM2.5-VL models for your specific use case to achieve the best results.

NotebookDescriptionLink
SFT (Unsloth)Supervised Fine-Tuning with LoRA using Unsloth.<a href="https://colab.research.google.com/drive/1FaR2HSe91YDe88TG97-JVxMygl-rL6vB?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHabLXysEu2E.png" width="110" alt="https://colab.research.google.com/github/Liquid4All/cookbook/blob/main/finetuning/notebooks/sftforvisionlanguagemodel.ipynb"></a>
SFT (TRL)Supervised Fine-Tuning with LoRA using TRL.<a href="https://colab.research.google.com/drive/10530jtJoa5zH2wgYlyXosypq1R7PIz?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHabLXysEu2E.png" width="110" alt="https://colab.research.google.com/github/Liquid4All/cookbook/blob/main/finetuning/notebooks/sftforvisionlanguagemodelwithtrl.ipynb"></a>

Performance

Benchmarks

LFM2.5-VL-3B significantly improves over LFM2-VL-3B in screen understanding, grounding, multi-image input, and tool use:

Benchmark**LFM2.5&#8209;VL&#8209;3B (3.1B)**LFM2&#8209;VL&#8209;3B (3.1B)Gemma4&nbsp;E2B (5.1B)Gemma4&nbsp;E4B (8B)InternVL&nbsp;3.5&nbsp;4B (4.7B)Qwen3.5-2B (2.3B)Qwen3.5-4B (4.7B)
ScreenSpot-v2 (avg)<ins>80.7</ins>-31.150.984.266.578.5
RefCOCO (Macro&nbsp;Prec@1)<ins>87.9</ins>57.1<br>(<span style="color: red;">-30.8</span>)67.372.188.978.586.6
BLINK<ins>61.5</ins>50.2<br>(<span style="color: red;">-11.3</span>)51.856.457.459.365.0
MuirBench<ins>58.3</ins>34.9<br>(<span style="color: red;">-23.4</span>)40.748.953.449.067.0
ToolSandBox59.526.4<br>(<span style="color: red;">-33.1</span>)56.5<ins>61.6</ins>n/a<a href="#fn-internvl-tools" class="footnote-ref" role="doc-noteref"><sup>1</sup></a>47.765.0
BFCLv432.520.5<br>(<span style="color: red;">-12.0</span>)33.2<ins>40.0</ins>n/a<a href="#fn-internvl-tools" class="footnote-ref" role="doc-noteref"><sup>1</sup></a>33.953.6

A selection of benchmarks for LFM2.5-VL-3B, including multimodal reasoning, math, and OCR (see our blog post for more benchmarks and details):

Benchmark**LFM2.5&#8209;VL&#8209;3B (3.1B)**LFM2&#8209;VL&#8209;3B (3.1B)Gemma4&nbsp;E2B (5.1B)Gemma4&nbsp;E4B (8B)InternVL&nbsp;3.5&nbsp;2B (2.4B)InternVL&nbsp;3.5&nbsp;4B (4.7B)Qwen3.5-2B (2.3B)Qwen3.5-4B (4.7B)
MME73.173.054.968.173.380.876.479.5
MMStar63.357.757.961.957.565.367.973.3
RealWorldQA73.171.156.261.861.468.671.476.2
CountBenchQA87.392.270.880.170.682.583.286.9
MMMB83.081.975.780.576.481.573.683.4
MM-IF Eval60.651.464.566.748.454.652.163.8
MathVista68.568.562.152.956.659.168.869.7
MMMU Pro30.528.734.539.127.431.643.560.9
ChartQA81.380.443.541.981.886.578.384.2
OCRBenchv2<a href="#fn-ocrbenchv2" class="footnote-ref" role="doc-noteref"><sup>2</sup></a>47.543.944.748.745.549.248.058.8
POPE88.789.284.086.987.388.988.786.0

<p id="fn-internvl-tools">[1]: InternVL 3.5 doesn't support tool use.</p>

<p id="fn-ocrbenchv2">[2]: English-only subset of OCRBenchv2.</p>

On-device Inference

LFM2.5-VL-3B decodes 228 tokens/s on an Apple M5 Max and 116 tokens/s on an AMD Ryzen AI Max+ 395, and fits in about 3 GB of memory. It even reaches 20 tokens/s on a Galaxy S26 Ultra, so you can run it fully on-device.

lfm2_5_vl_3b_on-device_inference_TTFT

GPU Inference

On a single NVIDIA H100 with vLLM, LFM2.5-VL-3B reaches the highest output throughput of any model we tested, about 11K tokens per second at high concurrency, or nearly 1B tokens per day.

lfm2_5_vl_3b_throughput

Because it answers directly instead of reasoning, LFM2.5-VL-3B is quick to first token on a single H100, reaching about 34 ms on a 5-frame clip.

lfm2_5_vl_3b_ttft

Contact

Citation

@article{liquidAI2026VL3B,
  author  = {Liquid AI},
  title   = {LFM2.5-VL-3B: A Better and Faster Vision-Language Model for the Edge},
  journal = {Liquid AI Blog},
  year    = {2026},
  note    = {www.liquid.ai/blog/lfm2-5-vl-3b},
}