Rithvik762/VLNCE-EnvDrop
VLNCE-EnvDrop Synthetic Vision-Language Navigation (VLN) data-augmentation set, derived from the EnvDrop augmentation used in VLN-CE / NaVILA-style training. Each of the 146,304 samples pairs a short first-person navigation video with the natural-language instruction the agent was following and the discrete action sequence it executed. This dataset provides the visual + motion supervision for training a GRU-augmented Qwen3-VL navigation model: the language conditions the… See the full description on the dataset page: https://huggingface.co/datasets/Rithvik762/VLNCE-EnvDrop.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face