CoolFace
Modelpublic

LiquidAI/LFM2-VL-450M

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
151likes14kdownloads
Model Card

<center> <div style="text-align: center;"> <img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png" alt="Liquid AI" style="width: 100%; max-width: 100%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;" /> </div> <div style="display: flex; justify-content: center; gap: 0.5em;"> <a href="https://playground.liquid.ai/chat?model=lfm2.5-vl-1.6b"><strong>Try LFM</strong></a> • <a href="https://docs.liquid.ai/lfm/getting-started/welcome"><strong>Docs</strong></a> • <a href="https://leap.liquid.ai/"><strong>LEAP</strong></a> • <a href="https://discord.com/invite/liquid-ai"><strong>Discord</strong></a> </div> </center>

<br>

LFM2‑VL-450M

LFM2‑VL is Liquid AI's first series of multimodal models, designed to process text and images with variable resolutions. Built on the LFM2 backbone, it is optimized for low-latency and edge AI applications.

We're releasing the weights of two post-trained checkpoints with 450M (for highly constrained devices) and 1.6B (more capable yet still lightweight) parameters.

  • 2× faster inference speed on GPUs compared to existing VLMs while maintaining competitive accuracy
  • Flexible architecture with user-tunable speed-quality tradeoffs at inference time
  • Native resolution processing up to 512×512 with intelligent patch-based handling for larger images, avoiding upscaling and distortion

Find more about our vision-language model in the LFM2-VL post and its language backbone in the LFM2 blog post.

📄 Model details

Due to their small size, we recommend fine-tuning LFM2-VL models on narrow use cases to maximize performance. They were trained for instruction following and lightweight agentic flows. Not intended for safety‑critical decisions.

Property[**LFM2-VL-450M**](https://huggingface.co/LiquidAI/LFM2-VL-450M)[**LFM2-VL-1.6B**](https://huggingface.co/LiquidAI/LFM2-VL-1.6B)
Parameters (LM only)350M1.2B
Vision encoderSigLIP2 NaFlex base (86M)SigLIP2 NaFlex shape‑optimized (400M)
Backbone layershybrid conv+attentionhybrid conv+attention
Context (text)32,768 tokens32,768 tokens
Image tokensdynamic, user‑tunabledynamic, user‑tunable
Vocab size65,53665,536
Precisionbfloat16bfloat16
LicenseLFM Open License v1.0LFM Open License v1.0

Supported languages: English

Generation parameters: We recommend the following parameters:

  • Text: temperature=0.1, min_p=0.15, repetition_penalty=1.05
  • Vision: min_image_tokens=64 max_image_tokens=256, do_image_splitting=True

Chat template: LFM2-VL uses a ChatML-like chat template as follows:

<|startoftext|><|im_start|>system
You are a helpful multimodal assistant by Liquid AI.<|im_end|>
<|im_start|>user
<image>Describe this image.<|im_end|>
<|im_start|>assistant
This image shows a Caenorhabditis elegans (C. elegans) nematode.<|im_end|>

Images are referenced with a sentinel (<image>), which is automatically replaced with the image tokens by the processor.

You can apply it using the dedicated `.apply_chat_template()` function from Hugging Face transformers.

Architecture

  • Hybrid backbone: Language model tower (LFM2-1.2B or LFM2-350M) paired with SigLIP2 NaFlex vision encoders (400M shape-optimized or 86M base variant)
  • Native resolution processing: Handles images up to 512×512 pixels without upscaling and preserves non-standard aspect ratios without distortion
  • Tiling strategy: Splits large images into non-overlapping 512×512 patches and includes thumbnail encoding for global context (in 1.6B model)
  • Efficient token mapping: 2-layer MLP connector with pixel unshuffle reduces image tokens (e.g., 256×384 image → 96 tokens, 1000×3000 → 1,020 tokens)
  • Inference-time flexibility: User-tunable maximum image tokens and patch count for speed/quality tradeoff without retraining

Training approach

  • Builds on the LFM2 base model with joint mid-training that fuses vision and language capabilities using a gradually adjusted text-to-image ratio
  • Applies joint SFT with emphasis on image understanding and vision tasks
  • Leverages large-scale open-source datasets combined with in-house synthetic vision data, selected for balanced task coverage
  • Follows a progressive training strategy: base model → joint mid-training → supervised fine-tuning

🏃 How to run LFM2-VL

You can run LFM2-VL with Hugging Face `transformers` v4.57 or more recent as follows:

bash
pip install -U transformers pillow

Here is an example of how to generate an answer with transformers in Python:

python
from transformers import AutoProcessor, AutoModelForImageTextToText
from transformers.image_utils import load_image

# Load model and processor
model_id = "LiquidAI/LFM2-VL-450M"
model = AutoModelForImageTextToText.from_pretrained(
    model_id,
    device_map="auto",
    dtype="bfloat16"
)
processor = AutoProcessor.from_pretrained(model_id)

# Load image and create conversation
url = "https://www.ilankelman.org/stopsigns/australia.jpg"
image = load_image(url)
conversation = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": image},
            {"type": "text", "text": "What is in this image?"},
        ],
    },
]

# Generate Answer
inputs = processor.apply_chat_template(
    conversation,
    add_generation_prompt=True,
    return_tensors="pt",
    return_dict=True,
    tokenize=True,
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=64)
processor.batch_decode(outputs, skip_special_tokens=True)[0]

# This image depicts a vibrant street scene in what appears to be a Chinatown or similar cultural area. The focal point is a large red stop sign with white lettering, mounted on a pole.

You can directly run and test the model with this Colab notebook.

🔧 How to fine-tune

We recommend fine-tuning LFM2-VL models on your use cases to maximize performance.

NotebookDescriptionLink
SFT (TRL)Supervised Fine-Tuning (SFT) notebook with a LoRA adapter using TRL.<a href="https://colab.research.google.com/drive/1csXCLwJx7wI7aruudBp6ZIcnqfv8EMYN?usp=sharing"><img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/vlOyMEjwHabLXysEu2E.png" width="110" alt="Colab link"></a>

📈 Performance

ModelRealWorldQAMM-IFEvalInfoVQA (Val)OCRBenchBLINKMMStarMMMU (Val)MathVistaSEEDBench_IMGMMVetMMEMMLU
InternVL3-2B65.1038.4966.1083153.1061.1048.7057.6075.0067.002186.4064.80
InternVL3-1B57.0031.1454.9479843.0052.3043.2046.9071.2058.701912.4049.80
SmolVLM2-2.2B57.5019.4237.7572542.3046.0041.6051.5071.3034.901792.50-
LFM2-VL-1.6B65.7546.3558.3572944.5049.8739.6751.7072.0047.751756.7250.99
ModelRealWorldQAMM-IFEvalInfoVQA (Val)OCRBenchBLINKMMStarMMMU (Val)MathVistaSEEDBench_IMGMMVetMMEMMLU
SmolVLM2-500M49.9011.2724.6460940.7038.2034.1037.5062.2029.901448.30-
LFM2-VL-450M52.0333.0944.5665742.6140.8734.4445.3063.6233.851229.9140.16

We obtained MM-IFEval and InfoVQA (Val) scores for InternVL 3 and SmolVLM2 models using VLMEvalKit.

📬 Contact

Citation

@article{liquidai2025lfm2,
 title={LFM2 Technical Report},
 author={Liquid AI},
 journal={arXiv preprint arXiv:2511.23404},
 year={2025}
}