koukeft/qwen3vl-2B-instruct-visqam-finetuned
Model Card for qwen3vl-2B-instruct-visqam-finetuned
Qwen3-VL-2B-Instruct fine-tuned on the VISQAM 1.2.0 dataset. This model is meant to improve the quality of prompt-based interactions with geographic maps.
Model Details
Model Description
- Developed by: Eftychia Koukouraki
- Funded by: NFDI4Earth
- Model type: Vision-Language Model (VLM) - Fine-tuned Qwen3-VL with LoRA
- Language(s) (NLP): English (en), Multilingual (inherits from Qwen3-VL)
- License: Apache 2.0
- Finetuned from model: Qwen/Qwen3-VL-2B-Instruct
Uses
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
Direct Use
- Visual question-answering on geographic maps
- Visual reasoning tasks for visualized geodata
Out-of-Scope Use
Not suitable for tasks outside the geospatial domain - performance may degrade significantly.
Bias, Risks, and Limitations
- This model inherits biases present in the base Qwen3-VL-2B-Instruct model, which was trained on large-scale internet data that may contain societal biases.
- Performance may vary across different geographic regions underrepresented in training data.
Recommendations
Always review model outputs before using them in production or decision-making contexts.
How to Get Started with the Model
The adapter model in this repository should be merged and unloaded. This can be done as following:
from transformers import AutoModelForVision2Seq
import torch
from peft import PeftModel
MODEL_NAME = 'Qwen/Qwen3-VL-2B-Instruct'
ADAPTER_PATH = 'path/to/downloaded/adapter/model'
base_model = AutoModelForVision2Seq.from_pretrained(
MODEL_NAME,
dtype="auto",
device_map="auto",
trust_remote_code=True
)
# Load adapter configuration and model
model = PeftModel.from_pretrained(base_model, ADAPTER_PATH)
# Optional: Merge adapter weights for faster inference
merged_model = model.merge_and_unload()
merged_model.save_pretrained("path/to/merged/model")
print("saved the merged model.")
The merged model can then be used like the Qwen3-VL-2B-Instruct base model:
from transformers import Qwen3VLForConditionalGeneration, AutoProcessor
import torch
inference_model = "path/to/fine-tuned/model"
# Load the model on the available device(s)
model = Qwen3VLForConditionalGeneration.from_pretrained(
inference_model, dtype="auto", device_map="auto"
)
processor = AutoProcessor.from_pretrained("Qwen/Qwen3-VL-2B-Instruct")
messages = [
{
"role": "user",
"content": [
{
"type": "image",
"image": "path/to/image.png",
},
{"type": "text", "text": "What is this map about?"},
],
}
]
# Preparation for inference
inputs = processor.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_dict=True,
return_tensors="pt"
)
inputs = inputs.to(model.device)
# Inference: Generation of the output
generated_ids = model.generate(**inputs, max_new_tokens=128)
generated_ids_trimmed = [
out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(output_text)
Training Details
Training Data
The training dataset for the fine-tuned model can be found at https://huggingface.co/datasets/koukeft/visqam/tree/1.2.0.
Training Procedure
The whole training procedure, as well as the split into 'train' and 'test' subsets is described in the following Python notebook: finetune_qwen3vl_on_visqam.ipynb.
Framework versions
- PEFT 0.17.1
Citation
If you use this model, please cite
@inproceedings{visqam_geoai2026,
author = {Koukouraki, Eftychia and Ajay, Ajay and Abubakar, Ahmad and Eid, Yomna},
title = {Introducing the VISQAM Dataset: Toward Automated Map Interpretation},
booktitle = {Proceedings of the 1st International Conference on Geospatial Artificial Intelligence (GeoAI 2026) – Oral Presentation Papers},
year = 2026,
publisher = {Zenodo},
address = {Ghent, Belgium},
doi = {10.5281/zenodo.20273245},
url = {https://doi.org/10.5281/zenodo.20273245},
}