Elliot89/Universal_Cross-Domain_Vision_Model
๐ฅ๐พ Universal Cross-Domain Vision Model
A multi-backbone vision model that classifies images across medical X-ray pathologies and sports action domains using fine-tuned multi-modal attention fusion on top of four pretrained encoders.
 
๐ง Model Architecture
The model fuses features from four pretrained backbone encoders through a learned multi-head attention fusion layer:
Each backbone's features are projected to a shared 512-dim space, then fused via an 8-head attention transformer block. The final classifier head outputs 14 class probabilities with an uncertainty estimate.
Image โ [BiomedCLIP, ViT-B/16, ResNet-50, EfficientNet-B0]
โ Projection Adapters (per backbone)
โ 8-Head Attention Fusion
โ Classifier โ 14 classes + Uncertainty estimate๐ท๏ธ Classes
๐ Running the Demo
Option 1 โ Hugging Face Spaces (live)
Visit the live demo โ no setup needed:
๐ https://huggingface.co/spaces/Elliot89/Universal_Cross-Domain_Vision_Model
Upload any image and click Classify.
Option 2 โ Run locally
Requirements: Python 3.9+, ~4 GB RAM (CPU) or GPU recommended
# 1. Clone this repo
git clone https://huggingface.co/spaces/Elliot89/Universal_Cross-Domain_Vision_Model
cd Universal_Cross-Domain_Vision_Model
# 2. Install dependencies
pip install -r requirements.txt
# 3. Launch
python app.py
# Opens at http://localhost:7860Option 3 โ REST API
# Start the API server
uvicorn api:app --host 0.0.0.0 --port 8000
# Classify an image file
curl -X POST http://localhost:8000/predict -F "file=@your_image.jpg"
# Classify from URL
curl -X POST http://localhost:8000/predict/url \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/xray.jpg"}'Interactive API docs at http://localhost:8000/docs
Option 4 โ Google Colab
Open colab_deploy.ipynb in Colab, set runtime to T4 GPU, and run all cells.
๐ฆ Repository Structure
โโโ app.py # Gradio web demo (main entry point)
โโโ api.py # FastAPI REST inference server
โโโ requirements.txt # Python dependencies
โโโ head_weights.pt # Fine-tuned fusion + classifier weights (~25 MB)
โโโ extract_head.py # Utility: extract head weights from full checkpoint
โโโ colab_deploy.ipynb # One-click Google Colab notebook
โโโ README.md # This fileNote on weights: The four backbone encoders (~1 GB total) are downloaded automatically from Hugging Face Hub at first startup and cached. Only the fine-tuned head (head_weights.pt, ~25 MB) is stored in this repo.๐ง Training Details
๐ API Response Format
{
"top_prediction": {
"label": "Pneumonia",
"confidence": 0.412
},
"predictions": [
{ "label": "Pneumonia", "confidence": 0.412 },
{ "label": "Normal", "confidence": 0.238 },
{ "label": "COVID-19", "confidence": 0.134 },
{ "label": "Tuberculosis", "confidence": 0.089 },
{ "label": "Cardiomegaly", "confidence": 0.061 },
{ "label": "Running", "confidence": 0.044 },
{ "label": "Lung Mass", "confidence": 0.031 },
{ "label": "Pleural Effusion","confidence": 0.021 }
]
}โ๏ธ Environment Variables
๐ ๏ธ Troubleshooting
Slow first startup โ The four backbones (~1 GB total) are downloaded from HF Hub on first run and cached. On HF Spaces this happens automatically during the build phase.
`head_weights.pt` not found โ The app still runs but uses random weights for the fusion and classifier layers. Predictions will not reflect actual training. Upload head_weights.pt to the repo to enable real predictions.
Out of memory โ The model runs on CPU if no GPU is detected. If memory is tight, reduce image resolution or comment out extra backbones in app.py.
Regenerating `head_weights.pt` from the original checkpoint โ If you have best_model_phase1.pt, run:
python extract_head.pyThis strips the large backbone weights (which are loaded from HF Hub) and saves only the fine-tuned layers (~25 MB) as head_weights.pt.
๐ License
MIT โ see https://opensource.org/licenses/MIT
๐ Acknowledgements
- Microsoft BiomedCLIP โ vision-language model pretrained on 15M medical image-text pairs from PubMed Central
- Stanford40 โ sports and human action recognition dataset
- timm โ PyTorch Image Models library
- open_clip โ open source CLIP implementation
- Gradio โ web demo framework
- FastAPI โ REST API framework
