prakhya15/trimodal-bind
0
TriModal-Bind
Retrieve the most semantically similar images from natural language using a shared multimodal embedding space learned through contrastive learning.
Features
- Text-to-image retrieval
- Contrastive multimodal embeddings
- DistilBERT text encoder
- PyTorch image encoder
- Gradio interface
