5ch4um1/lfm2.5-vrsbench-lora-450m
1255
LFM2.5-VL-450M VRSBench LoRA
Model Description
This is a fine-tuned version of LiquidAI's LFM2.5-VL-450M vision-language model, trained on the VRSBench dataset for general satellite image understanding. This serves as the base model for specialized satellite vision tasks.
The model can answer questions about satellite imagery, including:
- Scene classification
- Object detection and counting
- Visual question answering about satellite images
- General satellite image understanding
Training Details
VRSBench Training
- Base Model: LFM2.5-VL-450M
- Dataset: VRSBench (Vision Reasoning and Scene Understanding Benchmark)
- Epochs: 1
- Method: LoRA (r=16, alpha=32)
- Hardware: Local GPU training (no Ray/distributed)
Derived Models
This model serves as the base for specialized satellite vision experts:
Usage
With llama.cpp
# Download Q4_K_M quantized version (recommended)
wget https://huggingface.co/5ch4um1/lfm2.5-vrsbench-lora-450m/resolve/main/lfm2.5-vrsbench-lora-450m-q4_k_m.gguf
# Run inference
./llama-cli -m lfm2.5-vrsbench-lora-450m-q4_k_m.gguf \
--image satellite_image.jpg \
-p "Describe this satellite image in detail."With Transformers
from transformers import AutoModelForVision2Seq, AutoProcessor
from PIL import Image
model = AutoModelForVision2Seq.from_pretrained(
"5ch4um1/lfm2.5-vrsbench-lora-450m",
torch_dtype="auto",
device_map="auto"
)
processor = AutoProcessor.from_pretrained("5ch4um1/lfm2.5-vrsbench-lora-450m")
image = Image.open("satellite_image.jpg")
prompt = "What is shown in this satellite image?"
inputs = processor(text=prompt, images=image, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=100)
print(processor.decode(outputs[0], skip_special_tokens=True))GGUF Quantizations
Model Sources
- Base Model: LiquidAI/LFM2.5-VL-450M
- VRSBench Dataset: VRSBench Paper
Limitations
- General satellite understanding model - not specialized for specific tasks
- Performance varies depending on satellite image type and task
- For specialized tasks (terrain, maritime), use the derived expert models listed above
Training Environment
- Framework: Transformers + PEFT (LoRA)
- Hardware: Local GPU (CUDA)
- Training Scripts: Available in the cookbook repository
