CoolFace
Modelpublic

PeiyangLiu/CoE-Wiki-CoE-8B

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes16downloads
Model Card

CoE-Wiki-CoE-8B

CoE-Wiki-CoE-8B is an 8B vision-language checkpoint fine-tuned for Chain-of-Evidence question answering on Wiki-CoE. Given a question and candidate evidence screenshots, the model is trained to produce a structured answer with an evidence chain.

This checkpoint is intended for research on multimodal QA, visual evidence selection, and evidence-grounded reasoning over document-like screenshots.

Expected input and output

The model expects:

  • —a natural-language question
  • —candidate screenshot images that may contain the supporting evidence

The expected output is a JSON-style response with:

  • —evidence_chain: the selected supporting screenshots and localized evidence
  • —answer: the final answer

For exact prompt formatting and evaluation scripts, see the project code.

Usage

python
from transformers import AutoProcessor, AutoModelForImageTextToText
import torch

model_id = "PeiyangLiu/CoE-Wiki-CoE-8B"

processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True,
)

Use the same image preprocessing and prompt format as the CoE repository for reproducible results.

Related resources

  • —Homepage: https://lpy.pxsec.cn
  • —Paper: https://arxiv.org/abs/2605.01284
  • —Code: https://github.com/PeiYangLiu/CoE
  • —Dataset: https://huggingface.co/datasets/PeiyangLiu/wiki-coe
  • —SlideVQA checkpoint: https://huggingface.co/PeiyangLiu/CoE-SlideVQA-8B