Eku127/swiftvln-satnav-3b-1ep-f32s4-overlap0-pf-h8-pool-s2-posefilm
066
SwiftVLN for SatNav — Relative Pose
Related Repositories
- SwiftVLN: training and evaluation code for these checkpoints.
- SatNav: satellite-image navigation environments, datasets, and evaluation tools.
This checkpoint is designed for SwiftVLN on SatNav, the continuous-state vision-and-language navigation benchmark over satellite imagery. It augments visual tokens with relative pose through feature-wise linear modulation.
Model
- Starting checkpoint:
Qwen/Qwen2.5-VL-3B-Instruct - Training data: SatNav-v0.1 offline expert trajectories
- Training: 1 epoch, full-parameter fine-tuning, learning rate
2e-5 - Context: 32 RGB frames and up to 8 uniformly sampled history frames
- Prediction horizon: 4 actions
- Memory: per-frame average pooling with stride 2
- Pose fusion: FiLM over relative position and heading features
Repository name
swiftvln-satnav: SwiftVLN trained and evaluated on SatNav3b: Qwen2.5-VL 3B backbone1ep: trained for one epochf32s4: uses a 32-frame window and predicts four actionsoverlap0: uses non-overlapping training windowspf-h8-pool-s2: uses the reference per-frame memory configurationposefilm: injects relative pose into visual tokens with FiLM
Reference Results
Results reported in the SatNav paper, memory-design ablation, row Relative pose.
SR, SPL, and OS are percentages. Changes in SR and OS are percentage points relative to the SwiftVLN reference model.
Usage
Use this checkpoint with the SwiftVLN evaluation guide. Keep the full repository name unchanged because SwiftVLN derives the evaluation configuration from it.
This checkpoint is an input-augmentation ablation in the SwiftVLN SatNav Ablation Model Zoo.
