CoolFace
Modelpublic

koukeft/qwen3vl-2B-instruct-visqam-finetuned

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes13downloads
Model Card

Model Card for qwen3vl-2B-instruct-visqam-finetuned

Qwen3-VL-2B-Instruct fine-tuned on the VISQAM 1.2.0 dataset. This model is meant to improve the quality of prompt-based interactions with geographic maps.

Model Details

Model Description

  • —Developed by: Eftychia Koukouraki
  • —Funded by: NFDI4Earth
  • —Model type: Vision-Language Model (VLM) - Fine-tuned Qwen3-VL with LoRA
  • —Language(s) (NLP): English (en), Multilingual (inherits from Qwen3-VL)
  • —License: Apache 2.0
  • —Finetuned from model: Qwen/Qwen3-VL-2B-Instruct

Uses

<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->

Direct Use

  • —Visual question-answering on geographic maps
  • —Visual reasoning tasks for visualized geodata

Out-of-Scope Use

Not suitable for tasks outside the geospatial domain - performance may degrade significantly.

Bias, Risks, and Limitations

  • —This model inherits biases present in the base Qwen3-VL-2B-Instruct model, which was trained on large-scale internet data that may contain societal biases.
  • —Performance may vary across different geographic regions underrepresented in training data.

Recommendations

Always review model outputs before using them in production or decision-making contexts.

How to Get Started with the Model

The adapter model in this repository should be merged and unloaded. This can be done as following:

Python
from transformers import AutoModelForVision2Seq
import torch
from peft import PeftModel

MODEL_NAME = 'Qwen/Qwen3-VL-2B-Instruct'
ADAPTER_PATH = 'path/to/downloaded/adapter/model'

base_model = AutoModelForVision2Seq.from_pretrained(
    MODEL_NAME,
    dtype="auto",
    device_map="auto",
    trust_remote_code=True
)

# Load adapter configuration and model
model = PeftModel.from_pretrained(base_model, ADAPTER_PATH)

# Optional: Merge adapter weights for faster inference
merged_model = model.merge_and_unload()

merged_model.save_pretrained("path/to/merged/model")

print("saved the merged model.")

The merged model can then be used like the Qwen3-VL-2B-Instruct base model:

Python
from transformers import Qwen3VLForConditionalGeneration, AutoProcessor
import torch

inference_model = "path/to/fine-tuned/model"

# Load the model on the available device(s)
model = Qwen3VLForConditionalGeneration.from_pretrained(
    inference_model, dtype="auto", device_map="auto"
)

processor = AutoProcessor.from_pretrained("Qwen/Qwen3-VL-2B-Instruct")

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "image",
                "image": "path/to/image.png",
            },
            {"type": "text", "text": "What is this map about?"},
        ],
    }
]

# Preparation for inference
inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt"
)
inputs = inputs.to(model.device)

# Inference: Generation of the output
generated_ids = model.generate(**inputs, max_new_tokens=128)
generated_ids_trimmed = [
    out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
    generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(output_text)

Training Details

Training Data

The training dataset for the fine-tuned model can be found at https://huggingface.co/datasets/koukeft/visqam/tree/1.2.0.

Training Procedure

The whole training procedure, as well as the split into 'train' and 'test' subsets is described in the following Python notebook: finetune_qwen3vl_on_visqam.ipynb.

Framework versions

  • —PEFT 0.17.1

Citation

If you use this model, please cite

bibtex
@inproceedings{visqam_geoai2026,
  author = {Koukouraki, Eftychia and Ajay, Ajay and Abubakar, Ahmad and Eid, Yomna},
  title = {Introducing the VISQAM Dataset: Toward Automated Map Interpretation},
  booktitle = {Proceedings of the 1st International Conference on Geospatial Artificial Intelligence (GeoAI 2026) – Oral Presentation Papers},
  year = 2026,
  publisher = {Zenodo},
  address = {Ghent, Belgium},
  doi = {10.5281/zenodo.20273245},
  url = {https://doi.org/10.5281/zenodo.20273245},
}