CoolFace
Modelpublic

Ruggero1912/Patch-ioner_talk2dino_meacap_COCO_Captions

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes41downloads
Model Card

Patch-ionertalk2dinomeacapCOCOCaptions - Patch-ioner Configuration

This repository contains a pre-trained MEACAP model from the Patch-ioner framework for dense image captioning and controllable visual description.

๐Ÿ“ Paper Information

Title: "One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework" Authors: Lorenzo Bianchi, Giacomo Pacini, Fabio Carrara, Nicola Messina, Giuseppe Amato, Fabrizio Falchi ArXiv: https://arxiv.org/abs/2510.02898 Project Page: https://paciosoft.com/Patch-ioner/ Code: https://github.com/Ruggero1912/Patch-ioner

๐ŸŽฏ Model Overview

  • โ€”Model Type: MEACAP
  • โ€”Configuration: mlp.meacap.k.yaml
  • โ€”Vision Backbone: dinov2vitb14reg
  • โ€”Language Model: gpt2
  • โ€”Input Resolution: 518x518
  • โ€”Prefix Size: 768

MeaCap Configuration

  • โ€”Project Length: 10
  • โ€”Temperature: 0.01
  • โ€”Top-K: 3
  • โ€”Memory Caption Num: 5
  • โ€”VL Model: openai/clip-vit-base-patch16
  • โ€”WTE Model: sentence-transformers/all-MiniLM-L6-v2
  • โ€”Parser Checkpoint: lizhuang144/flan-t5-base-VG-factual-sg
  • โ€”Memory ID: cocoB16t2d
  • โ€”Entity Retrieval: coco_entities

๐Ÿ“Š Performance

TaskMETEORCIDErSPICE
Image Captioning0.2070.7170.157
Narratives10.00027.40012.700

๐Ÿ“ˆ Detailed Results

Image Captioning Results

  • โ€”METEOR: 0.2075
  • โ€”CIDEr: 0.7175
  • โ€”SPICE: 0.1573
  • โ€”BLEU_4: 0.1968
  • โ€”ROUGE_L: 0.4200
  • โ€”CLIP-S: 0.7278

Narratives Results

  • โ€”METEOR: 10.0000
  • โ€”CIDEr: 27.4000
  • โ€”SPICE: 12.7000
  • โ€”BLEU_4: 2.4000
  • โ€”ROUGE_L: 20.2000
  • โ€”CLIP-S: 67.4000

๐Ÿš€ Quick Start

python
from transformers import AutoModel
import torch
from PIL import Image

MODEL_ID = "Ruggero1912/Patch-ioner_talk2dino_meacap_COCO_Captions"

# Load the model with AutoModel from the transformers library
model = AutoModel.from_pretrained(MODEL_ID, trust_remote_code=True)

# Example image (replace with your actual image loading logic)
# For a real scenario, you would load an image from a file or URL.
# e.g., image = Image.open("path/to/your/image.jpg")
image = Image.new('RGB', (224, 224), color = 'red') # Placeholder image

# The specific `forward` method signature depends on the model's implementation
# within the `patchioner` library. You might need to preprocess the image
# and provide additional inputs (e.g., text prompts for controllable captioning).
# Please refer to the official GitHub repository for detailed inference examples
# using the `Patchioner` library's specific `forward` methods.

# If the model has a simplified call for basic captioning, it might look like this:
# results = model(image)
# print(results)
print(f"Model {MODEL_ID} loaded successfully using `transformers.AutoModel`. "
      "Refer to the original Patch-ioner GitHub for full usage details and example inference.")

๐Ÿ“ Repository Contents

  • โ€”config.yaml: Model configuration file
  • โ€”model.pt: Pre-trained model weights
  • โ€”memory_captions.json: MeaCap memory captions database
  • โ€”memory_clip_embeddings.pt: MeaCap CLIP embeddings for memory
  • โ€”memory_wte_embeddings.pt: MeaCap WTE embeddings for memory- README.md: This file

๐Ÿ”ง Installation

bash
pip install git+https://github.com/Ruggero1912/Patch-ioner

๐Ÿ’ก Usage Examples

Refer to the Patch-ioner repository for updated usage examples.

๐ŸŽ›๏ธ Model Configuration

  • โ€”Prefix Size: 768
  • โ€”Memory Bank Size: 0
  • โ€”Normalization: False

๐Ÿ“ˆ Training Details

  • โ€”Training Dataset: COCO Captions
  • โ€”Training Epochs: TBD
  • โ€”Batch Size: TBD
  • โ€”Learning Rate: TBD
  • โ€”Optimizer: AdamW

๐Ÿ“š Citation

If you use this model in your research, please cite our paper, refer to the Project Page for updated citation template.

๐Ÿค Contributing

We welcome contributions to improve the Patch-ioner framework. Please see the main repository for contribution guidelines.

๐Ÿ“„ License

See the main repository for detailed license information.

๐Ÿ› Issues and Support

For issues related to this model or the Patch-ioner framework, please:

  1. 1.Check the main repository for existing issues
  2. 2.Open a new issue with detailed information about your problem
  3. 3.Contact the authors.

๐Ÿ”— Related Models

Explore other Patch-ioner model configurations:

More models available in [Ruggero1912's models](https://huggingface.co/Ruggero1912)


This model is part of the Patch-ioner framework for dense image captioning and controllable visual description.