maelic/relsgg-vits16plus
relsgg-vits16plus
Open-vocabulary relation prediction from any boxes or masks. Give the model an image and regions from any source (a detector, a segmenter, ground truth); it returns ranked relations over a predicate vocabulary supplied at inference, and optionally two graphs (spatial + semantic) from the same forward pass. Object class labels are never an input.
Part of RelateAnything (code · paper). Trained on RA-4M; evaluated with OV-SGG-Bench.
Use it
pip install git+https://github.com/Maelic/RelateAnything
hf download maelic/relsgg-vits16plus # optional; the API fetches on first usefrom relsgg import RelateAnything
# Regions come from any detector, any segmenter, or your own annotation.
# Object class labels are never an input.
model = RelateAnything.from_pretrained("maelic/relsgg-vits16plus", device="cuda")
for t in model.predict(image, boxes_xyxy, topk=20): # PIL/ndarray, boxes [N, 4] in pixels
print(t) # (person) --riding [0.67]--> (horse)
# Masks instead of boxes: pass the [N, H, W] binary masks beside their extents.
triplets = model.predict(image, boxes_xyxy, masks=masks, topk=20)
# The vocabulary is an input. Any strings, at any time, without retraining.
model.set_vocabulary(["about to collide with", "reflected in"])
# Or answer from the whole training vocabulary, 19,103 strings, read from the weights.
model = RelateAnything.from_pretrained("maelic/relsgg-vits16plus", full_vocabulary=True, device="cuda")
# Two graphs from one forward pass.
graphs = model.predict(image, boxes_xyxy, decompose=True) # {"spatial": [...], "semantic": [...]}Every vocabulary is encoded once by the text student shipped beside the weights, and the head is reparameterized onto it; scoring afterwards is vision only. full_vocabulary=True reads predicate_embeddings.npz instead of encoding, which turns a minute and a half of CPU work into a download. model.pth embeds the backbone configuration, so running these weights needs no gated DINOv3 login.
Files: model.pth (torch, EMA weights), text_student.pt + tokenizer, predicate_embeddings.npz (the training vocabulary, encoded), relateanything.onnx, predicate_bank.npz, thresholds.json, calibration.json, README.md.
Every number below is generated from measured eval artifacts (`release/make_model_cards.py`); none is hand-typed.
Closed-vocabulary transfer (reparameterized, TEST, graph-constrained)
Open-vocabulary, NO reparameterization (all 19,103 predicates deployed)
Synonym-matched at the calibrated tau (see provenance). This is the honest "the model never saw your label set" protocol.
Spatial reasoning (SpatialSense, adversarial true/false; chance = 0.5)
Macro AUC over predicates: 0.6897
Two-graph decomposition (spatial / semantic, type-stratified protocol)
Deployment thresholds (per-predicate best-F1, measured on THIS checkpoint)
Score scales are checkpoint-specific (the output head is rank-trained), so these thresholds transfer to no other model. Regime: gt boxes, pair_weight=0, 5000 val images. Top predicates by support:
Provenance
License and data notices
Weights are a derivative of Meta DINOv3 pretrained weights and are distributed under the DINOv3 license. Training annotations (RA-4M) were generated by gemma-4-26B and carry the Gemma Terms of Use notice; images are referenced by identifier only (Objects365/COCO/OpenImages). The vg_raw subset derives from Visual Genome (CC BY 4.0). Predicate synonyms are deliberately never collapsed — surface-form diversity is part of the label space. Full notices: THIRD_PARTY_NOTICES.md in the code repository.
Citation
@article{neau2026relateanything,
title = {RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs},
author = {Neau, Ma\"elic},
journal = {arXiv preprint arXiv:2609.12552},
eprint = {2609.12552},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2609.12552},
year = {2026}
}