CoolFace
Modelpublic

Eku127/swiftvln-satnav-3b-1ep-f32s4-overlap0-pf-h8-pool-s2-posefilm

sourceHugging Faceupdated 9d agoView on Hugging Face
0likes66downloads
Model Card

SwiftVLN for SatNav — Relative Pose

Related Repositories

  • —SwiftVLN: training and evaluation code for these checkpoints.
  • —SatNav: satellite-image navigation environments, datasets, and evaluation tools.

This checkpoint is designed for SwiftVLN on SatNav, the continuous-state vision-and-language navigation benchmark over satellite imagery. It augments visual tokens with relative pose through feature-wise linear modulation.

Model

  • —Starting checkpoint: Qwen/Qwen2.5-VL-3B-Instruct
  • —Training data: SatNav-v0.1 offline expert trajectories
  • —Training: 1 epoch, full-parameter fine-tuning, learning rate 2e-5
  • —Context: 32 RGB frames and up to 8 uniformly sampled history frames
  • —Prediction horizon: 4 actions
  • —Memory: per-frame average pooling with stride 2
  • —Pose fusion: FiLM over relative position and heading features

Repository name

  • —swiftvln-satnav: SwiftVLN trained and evaluated on SatNav
  • —3b: Qwen2.5-VL 3B backbone
  • —1ep: trained for one epoch
  • —f32s4: uses a 32-frame window and predicts four actions
  • —overlap0: uses non-overlapping training windows
  • —pf-h8-pool-s2: uses the reference per-frame memory configuration
  • —posefilm: injects relative pose into visual tokens with FiLM

Reference Results

Results reported in the SatNav paper, memory-design ablation, row Relative pose.

SR, SPL, and OS are percentages. Changes in SR and OS are percentage points relative to the SwiftVLN reference model.

SplitSR (%)SPL (%)OS (%)Delta SR (pp)Delta OS (pp)
Test Seen (val_seen)67.867.576.3+2.0+3.9
Test Unseen (val_unseen)55.455.265.0+1.7+0.9

Usage

Use this checkpoint with the SwiftVLN evaluation guide. Keep the full repository name unchanged because SwiftVLN derives the evaluation configuration from it.

This checkpoint is an input-augmentation ablation in the SwiftVLN SatNav Ablation Model Zoo.