princetunes/gold-weight-prediction-qwen2.5-vl
Qwen2.5-VL-3B-Instruct Gold Weight Predictor (LoRA)
An AI model fine-tuned for predicting the exact gram weight of gold jewelry designs from single images. Developed specifically to address the "Abstract Unit Trap" where standard multimodal models fail to interpret thickness, solid vs. hollow shanks, setting volume, and density cohesion.
It is fine-tuned on a unified, high-quality catalog dataset derived from Orra and Tanishq jewelry collections (covering 14KT, 18KT, and 22KT gold rings across Indian and US size standards).
๐ Performance Comparison
<p align="center"> <img src="mae_comparison.png" width="85%" alt="Gold Weight Prediction MAE and MAPE Comparison" /> </p>
<p align="center"> <img src="predictions_scatter.png" width="65%" alt="Gold Weight Prediction Correlation Scatter Plot" /> </p>
Our fine-tuned Qwen2.5-VL model (evaluated in a native direct-baseline configuration) is compared here against Gemini 3.1 Flash Lite Preview (running with a production-grade 5-item Visual Retrieval-Augmented Generation (RAG) pipeline) and the base model on validation split examples.
Key Metrics
[!IMPORTANT] Key Finding: Fine-tuning the vision-language model directly on gold jewelry weights resulted in a 25.9% reduction in Mean Absolute Error (MAE) compared to the production-grade Gemini 3.1 Flash RAG system, while completely eliminating the need for vector database queries and CLIP embedding overhead during inference.
๐ Why It Outperforms Generalist Models
- Understanding Solid vs. Hollow Geometry: The model has been trained to recognize structural cues (shank thickness, setting volume, under-gallery voids) from images to estimate volume.
- Purity and Size Density Cohesion: Standard LLMs struggle with physical formulas. This model learned to calibrate predictions directly based on the structural volume of the design, the targeted metal purity (14K, 18K, 22K), and the target ring size.
- Calibrated Scale Anchors: Generalist VLM weights tend to be random guesses. The fine-tuned Qwen2.5-VL maps visual shapes to real gold weights.
๐ ๏ธ How to Use
Loading the Model & Processors
import torch
from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor
from peft import PeftModel
# Base model and adapter paths
base_model_name = "Qwen/Qwen2.5-VL-3B-Instruct"
adapter_model_path = "princetunes/gold-weight-prediction-qwen2.5-vl"
# 1. Load the processor and base model
processor = AutoProcessor.from_pretrained(base_model_name)
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
base_model_name,
torch_dtype=torch.bfloat16,
device_map="auto"
)
# 2. Attach the fine-tuned LoRA adapters
model = PeftModel.from_pretrained(model, adapter_model_path)Performing Inference
from PIL import Image
import requests
def predict_gold_weight(image_path, purity="18K", ring_size=6.0):
image = Image.open(image_path).convert("RGB")
# Prompt format used during fine-tuning
prompt = f"Analyze this jewelry ring design. Predict its 18K gold weight in grams for ring size {ring_size}."
# Prepare inputs using the standard Qwen2.5-VL format
messages = [
{
"role": "user",
"content": [
{"type": "image", "image": image},
{"type": "text", "text": prompt}
]
}
]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, video_inputs = processor.image_processor(images=image, videos=None, return_tensors="pt")
inputs = processor(
text=[text],
images=image,
padding=True,
return_tensors="pt"
).to(model.device)
# Generate weight prediction
with torch.no_grad():
generated_ids = model.generate(**inputs, max_new_tokens=32)
generated_ids_trimmed = [
out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)[0]
return output_text
# Example execution
print(predict_gold_weight("ring_image.jpg", purity="18K", ring_size=7.0))โ๏ธ Fine-Tuning Specifications
- Base Model: Qwen2.5-VL-3B-Instruct
- Parameters Trained: LoRA adapters (Rank
r = 16, Alphalora_alpha = 32, Target Modules: Causal LM projections) - Dropout: 0.05
- Epochs: 3
- Dataset: 1,600+ jewelry designs (Orra + Tanishq unified dataset) mapped with multiple visual angles and precise density metadata.
