CoolFace
Modelpublic

princetunes/gold-weight-prediction-qwen2.5-vl

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes4downloads
Model Card

Qwen2.5-VL-3B-Instruct Gold Weight Predictor (LoRA)

An AI model fine-tuned for predicting the exact gram weight of gold jewelry designs from single images. Developed specifically to address the "Abstract Unit Trap" where standard multimodal models fail to interpret thickness, solid vs. hollow shanks, setting volume, and density cohesion.

It is fine-tuned on a unified, high-quality catalog dataset derived from Orra and Tanishq jewelry collections (covering 14KT, 18KT, and 22KT gold rings across Indian and US size standards).


๐Ÿ“Š Performance Comparison

<p align="center"> <img src="mae_comparison.png" width="85%" alt="Gold Weight Prediction MAE and MAPE Comparison" /> </p>

<p align="center"> <img src="predictions_scatter.png" width="65%" alt="Gold Weight Prediction Correlation Scatter Plot" /> </p>

Our fine-tuned Qwen2.5-VL model (evaluated in a native direct-baseline configuration) is compared here against Gemini 3.1 Flash Lite Preview (running with a production-grade 5-item Visual Retrieval-Augmented Generation (RAG) pipeline) and the base model on validation split examples.

Key Metrics

Model ConfigurationEvaluation TypeMean Absolute Error (MAE)Mean Absolute Percentage Error (MAPE)Key Advantage
Qwen2.5-VL Fine-Tuned (Direct)Native Inference0.718 g37.16%Best Performance (No external DB queries required)
Qwen2.5-VL Fine-Tuned (RAG)With Visual Anchors1.005 g45.06%High consistency but slightly higher variance
Gemini 3.1 Flash (RAG Avg)Center-of-Range RAG0.969 g46.32%Requires expensive embedding & Pinecone query
Gemini 3.1 Flash (RAG Max)Production RAG (Max)1.062 g51.42%High bias towards upper-limit estimations
Qwen2.5-VL Base (Un-tuned)Native Prompting> 3.500 g> 150.0%Hallucinates abstract volumes; no density cohesion
[!IMPORTANT] Key Finding: Fine-tuning the vision-language model directly on gold jewelry weights resulted in a 25.9% reduction in Mean Absolute Error (MAE) compared to the production-grade Gemini 3.1 Flash RAG system, while completely eliminating the need for vector database queries and CLIP embedding overhead during inference.

๐Ÿ” Why It Outperforms Generalist Models

  1. 1.Understanding Solid vs. Hollow Geometry: The model has been trained to recognize structural cues (shank thickness, setting volume, under-gallery voids) from images to estimate volume.
  2. 2.Purity and Size Density Cohesion: Standard LLMs struggle with physical formulas. This model learned to calibrate predictions directly based on the structural volume of the design, the targeted metal purity (14K, 18K, 22K), and the target ring size.
  3. 3.Calibrated Scale Anchors: Generalist VLM weights tend to be random guesses. The fine-tuned Qwen2.5-VL maps visual shapes to real gold weights.

๐Ÿ› ๏ธ How to Use

Loading the Model & Processors

python
import torch
from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor
from peft import PeftModel

# Base model and adapter paths
base_model_name = "Qwen/Qwen2.5-VL-3B-Instruct"
adapter_model_path = "princetunes/gold-weight-prediction-qwen2.5-vl"

# 1. Load the processor and base model
processor = AutoProcessor.from_pretrained(base_model_name)
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
    base_model_name,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

# 2. Attach the fine-tuned LoRA adapters
model = PeftModel.from_pretrained(model, adapter_model_path)

Performing Inference

python
from PIL import Image
import requests

def predict_gold_weight(image_path, purity="18K", ring_size=6.0):
    image = Image.open(image_path).convert("RGB")
    
    # Prompt format used during fine-tuning
    prompt = f"Analyze this jewelry ring design. Predict its 18K gold weight in grams for ring size {ring_size}."
    
    # Prepare inputs using the standard Qwen2.5-VL format
    messages = [
        {
            "role": "user",
            "content": [
                {"type": "image", "image": image},
                {"type": "text", "text": prompt}
            ]
        }
    ]
    
    text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
    image_inputs, video_inputs = processor.image_processor(images=image, videos=None, return_tensors="pt")
    
    inputs = processor(
        text=[text],
        images=image,
        padding=True,
        return_tensors="pt"
    ).to(model.device)
    
    # Generate weight prediction
    with torch.no_grad():
        generated_ids = model.generate(**inputs, max_new_tokens=32)
        generated_ids_trimmed = [
            out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
        ]
        output_text = processor.batch_decode(
            generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
        )[0]
        
    return output_text

# Example execution
print(predict_gold_weight("ring_image.jpg", purity="18K", ring_size=7.0))

โš™๏ธ Fine-Tuning Specifications

  • โ€”Base Model: Qwen2.5-VL-3B-Instruct
  • โ€”Parameters Trained: LoRA adapters (Rank r = 16, Alpha lora_alpha = 32, Target Modules: Causal LM projections)
  • โ€”Dropout: 0.05
  • โ€”Epochs: 3
  • โ€”Dataset: 1,600+ jewelry designs (Orra + Tanishq unified dataset) mapped with multiple visual angles and precise density metadata.