E27085921/HIKARI-Vega-8B-SkinCaption-Fused-LoRA
<p align="center"> <img src="HIKARI_logo.png" alt="HIKARI" width="100%"/> </p>
<h1 align="center">HIKARI-Vega-8B-SkinCaption-Fused-LoRA</h1>
<p align="center"> <img src="https://img.shields.io/badge/Type-LoRA%20Adapter-blueviolet?style=flat-square"/> <img src="https://img.shields.io/badge/Size-~1.1%20GB-lightblue?style=flat-square"/> <img src="https://img.shields.io/badge/Base-Qwen3--VL--8B--Thinking-blue?style=flat-square"/> <img src="https://img.shields.io/badge/License-Apache%202.0-orange?style=flat-square"/> </p>
๐ Model Type: LoRA Adapter
This is a LoRA adapter (~1.1 GB) โ it must be loaded on top of the merged Stage 2 modelE27085921/HIKARI-Sirius-8B-SkinDx-RAG. โ ๏ธ Important: Do NOT load this on top of rawQwen/Qwen3-VL-8B-Thinking. The merged-init design requires disease knowledge to already be in the base weights. UseHIKARI-Sirius-8B-SkinDx-RAG(merged) as the base. โ Advantage: Lightweight โ download only ~1.1 GB instead of ~17 GB (plus the ~17 GB Stage 2 base). ๐พ If you prefer a standalone ready-to-use model, see the merged version: [E27085921/HIKARI-Vega-8B-SkinCaption-Fused](https://huggingface.co/E27085921/HIKARI-Vega-8B-SkinCaption-Fused) (~17 GB)
What is this adapter?
LoRA adapter for [HIKARI-Vega-8B-SkinCaption-Fused](https://huggingface.co/E27085921/HIKARI-Vega-8B-SkinCaption-Fused) โ Clinical skin lesion caption generation with merged-init strategy (best caption model). Metric: BLEU-4: 29.33 โญ.
This is the LoRA adapter for the best caption generation model in the HIKARI family. Note: for correct behavior, load this adapter on top of the merged Stage 2 model (HIKARI-Sirius-8B-SkinDx-RAG), not the raw base model โ disease knowledge must already be in the base weights.
See the full model card at [E27085921/HIKARI-Vega-8B-SkinCaption-Fused](https://huggingface.co/E27085921/HIKARI-Vega-8B-SkinCaption-Fused) for complete details, usage examples, and performance comparison.
Usage
from peft import PeftModel
from transformers import Qwen3VLForConditionalGeneration, AutoProcessor
import torch
from PIL import Image
# Step 1: Load the MERGED Stage 2 disease model as base
# (disease knowledge is permanently in its weights โ DO NOT use raw Qwen base here)
base = Qwen3VLForConditionalGeneration.from_pretrained(
"E27085921/HIKARI-Sirius-8B-SkinDx-RAG", # merged Stage 2 weights (~17 GB)
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True,
)
# Step 2: Apply Stage 3 caption LoRA adapter (~1.1 GB)
model = PeftModel.from_pretrained(base, "E27085921/HIKARI-Vega-8B-SkinCaption-Fused-LoRA")
processor = AutoProcessor.from_pretrained("E27085921/HIKARI-Vega-8B-SkinCaption-Fused-LoRA", trust_remote_code=True)
# Step 3: Inference โ see full examples at E27085921/HIKARI-Vega-8B-SkinCaption-Fused
image = Image.open("skin_lesion.jpg").convert("RGB")For complete inference examples including vLLM and SGLang production code, see: [E27085921/HIKARI-Vega-8B-SkinCaption-Fused](https://huggingface.co/E27085921/HIKARI-Vega-8B-SkinCaption-Fused)
๐ Citation
@misc{hikari2026,
title = {HIKARI: RAG-in-Training for Skin Disease Diagnosis
with Cascaded Vision-Language Models},
author = {Watin Promfiy and Pawitra Boonprasart},
year = {2026},
institution = {King Mongkut's Institute of Technology Ladkrabang,
Department of Information Technology, Bangkok, Thailand}
}<p align="center">Made with โค๏ธ at <b>King Mongkut's Institute of Technology Ladkrabang (KMITL)</b></p>
