Voxel51/openworld-sam
1
OpenWorldSAM — Zero-Shot Universal Image Segmentation
OpenWorldSAM extends SAM2 with language understanding to enable open-vocabulary instance, panoptic, semantic, and referring segmentation from arbitrary text prompts — without any task-specific fine-tuning.
This repository is an unofficial HuggingFace Hub mirror of the GinnyXiao/OpenWorldSAM checkpoint, uploaded to enable loading via transformers and the FiftyOne Model Zoo. All credit belongs to the original authors.
Attribution
Paper: "Extending SAM2 for Universal Image Segmentation with Language Prompts" Fangxun Shu\, Yiwen Ye\, Jianhua Han, Jiwen Yu, Qize Yang, Xiao-Ping Zhang, Hang Xu, Bei Yu, Xiaodan Liang. NeurIPS 2025 Spotlight · arXiv 2507.05427
Original code: GinnyXiao/OpenWorldSAM — Apache-2.0 License
Checkpoint: ADE20K instance segmentation (openworld_sam_ade20k.pt) from the official Google Drive releaseUsage
Via FiftyOne Model Zoo
import fiftyone as fo
import fiftyone.zoo as foz
dataset = foz.load_zoo_dataset("quickstart", max_samples=5)
model = foz.load_zoo_model(
"openworld-sam-ade20k-torch",
class_names=["person", "car", "chair", "table", "sky", "tree"],
)
dataset.apply_model(model, label_field="owsam_pred")
session = fo.launch_app(dataset)Standalone (trust_remote_code)
import torch
from transformers import AutoModel
model = AutoModel.from_pretrained(
"neerajaabhyankar/openworld-sam", trust_remote_code=True
)
model.eval()
import numpy as np
arr = np.array(your_pil_image) # HWC uint8 RGB
batched_inputs = [{
"image": model.preprocess_image(arr),
"evf_image": model.preprocess_image_beit3(arr),
"height": arr.shape[0],
"width": arr.shape[1],
"prompt": ["person", "car", "tree"],
"unique_categories": [0, 1, 2],
}]
with torch.no_grad():
outputs = model(batched_inputs)
# outputs[0]["instances"]:
# "masks" — bool tensor [N, H, W]
# "scores" — float tensor [N]
# "class_ids" — long tensor [N]Architecture
Notes
Loading this checkpoint via from_pretrained prints a LOAD REPORT with a few MISSING/UNEXPECTED keys. Both are understood and safe to ignore for the single-image zero-shot segmentation path documented above:
- *`evf_sam2.text_hidden_fcs.
— MISSING (dead duplicate module, harmless).** BothEvfSam2Model._init` (`model/evfsam2.py) andOpenWorldSAMModel._init` (`modelingopenworldsam.py`) independently construct their own `texthiddenfcs` projection `ModuleList`. The checkpoint only ever stored one copy, at the top-level `texthiddenfcs.*` key — that's the one `OpenWorldSAMModel` actually uses in `forward()` (`self.texthiddenfcs[0](feat)`) and it loads correctly. The nested `evfsam2.texthiddenfcscopy is never referenced anywhere inOpenWorldSAMModel.forward(), so it gets randomly initialized and simply sits unused. No effect on inference output. (Could be cleaned up by removing the unused construction inEvfSam2Model._init_`, but there's no correctness reason to.)
- `memory_encoder.fuser.layers.{0,1}` — gamma/weight name mismatch (real bug, but on an unused code path). In
model/segment_anything_2/sam2/modeling/memory_encoder.py,CXBlockhas a comment claiming the layer-scale parameter was renamed:# modified by ZhangYx from self.gamma to self.weight. Due to (facebookresearch/segment-anything-2#85)— but the code still declaresself.gamma. The checkpoint was produced by the actual ZhangYx fork, which did apply the rename, so its keys are...weight. Since the bundled code never got the rename,gammaandweightdon't match: the checkpoint'sweightvalues show as UNEXPECTED, and the model'sgammaparameters show as MISSING (randomly initialized instead of loaded). In practice this doesn't affect the model as used here:memory_encoder(and itsfuser) is only invoked from SAM2's video mask-propagation path (self.memory_encoder(...)insam2_base.py, called during track-step), whichOpenWorldSAMModel.forward()never exercises — it only callsvisual_model.forward_image()+_prepare_backbone_features()for single-image inference. This would only matter if a future integration adds video/memory-based mask propagation on top of this checkpoint; at that point, fix it by renamingself.gamma→self.weightinCXBlockto match the checkpoint.
Requirements
torch torchvision transformers safetensors timm einopsNo detectron2 required — this mirror is self-contained.
License
Apache-2.0 (same as original GinnyXiao/OpenWorldSAM)
