datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ui-navigation-corpus
User Interface (Navigation) Corpus
Overview
This dataset serves as a collection of various images of, videos and metadata of mobile (both iOS and Android) and web user interfaces as well as tags and text extractions associated to them.
Dataset also includes user interface navigation annotations and videos related to them. One of the possible use cases of this dataset is training a UI navigation agent.
Dataset Structure
The resources of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/teleren/ui-navigation-corpus.ui-navigation-corpus
User Interface (Navigation) Corpus
Overview
This dataset serves as a collection of various images of, videos and metadata of mobile (both iOS and Android) and web user interfaces as well as tags and text extractions associated to them.
Dataset also includes user interface navigation annotations and videos related to them. One of the possible use cases of this dataset is training a UI navigation agent.
Dataset Structure
The resources of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/ijlewis/ui-navigation-corpus.labutopia_level5_navigation_20260826navigation_aid_2construction-site-robot-navigation-vla-training-next-pack-221fb1bd-ce8fc779
Aerial Outdoor Scenes for Drone Ice-on-Power-Line Detection (YOLO)
Training dataset for a drone-mounted YOLO model that detects ice accretion on power lines during aerial inspection. Renders are staged in outdoor open-air scenes at 1280x1280 with RGB, albedo and world-space normal (OpenGL, linear) passes, per-frame annotations (bounding boxes, camera pose, labels), and authored daylight. Note: the available environments do not frame actual power lines, so the set stands in for… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/construction-site-robot-navigation-vla-training-next-pack-221fb1bd-ce8fc779.sketchvlm-maze-navigation
SketchVLM: Maze Navigation
This dataset is associated with the paper: SketchVLM: Vision Language Models Can Annotate Images to Explain Thoughts and Guide Users.
SketchVLM is a training-free, model-agnostic framework that enables Vision-Language Models (VLMs) to produce non-destructive, editable SVG overlays on input images to visually explain their answers. The Maze Navigation dataset is one of the benchmarks introduced to evaluate a model's ability to trace a path from start to end… See the full description on the dataset page: https://huggingface.co/datasets/loganbolton/sketchvlm-maze-navigation.Spatial_Navigation
🌟 This repo contains part of the training dataset for model ThinkMorph-7B.
Dataset Description
We create an enriched interleaved dataset centered on four representative tasks requiring varying degrees of visual engagement and cross-modal interactions, including Jigsaw Assembly, Spatial Navigation, Visual Search and Chart Refocus.
Statistics
Dataset Usage
Data Downloading… See the full description on the dataset page: https://huggingface.co/datasets/ThinkMorph/Spatial_Navigation.openapps-scripted-navigation-3024ep
OpenApps Scripted Navigation (3,024 episodes, 20 routes)
Inter-app navigation trajectories in OpenApps, collected with a scripted
policy using only real UI actions (click "Return to List of Apps", click the
target app icon) — no goto() teleports, so transitions are learnable and
plannable from the action space alone.
20 routes: 5 source apps (todo, calendar, messages, codeeditor, map) × 4 targets
3,024 episodes after filtering (from 4,000 collected), 20 steps each
Episode… See the full description on the dataset page: https://huggingface.co/datasets/FruitPunchSamuraiG/openapps-scripted-navigation-3024ep.vamos_navigation_only_dataset
vamos_navigation_only_dataset
Description
VAMOS navigation-only (without natural language preference annotations nor co-training VQA data) dataset with 100% of TartanDrive 2 data, 50% of SCAND data, 25% of CODa data, and 100% of in-domain spot data. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5.
License
This dataset is released under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vamos_navigation_only_dataset.InVLM-navigation-dataset2Simaihub_Mobile_Robot_Navigation_Training_Dataset_sampleSimaihub Mobile Robot Navigation Training Dataset (Isaac Lab)
Version: V2.0 , Release Date: 2026-08-14
Short Blurb
Licensed expert navigation trajectories for offline RL and imitation learning — multi-scene indoor AMR data with domain randomization, dynamic obstacles,
and recovery behaviors.
In addition to core navigation datasets, we will continue to expand simulation datasets for segmented embodied intelligence scenarios, including industrial robotic arm operation, warehouse… See the full description on the dataset page: https://huggingface.co/datasets/Simaihub/Simaihub_Mobile_Robot_Navigation_Training_Dataset_sample.InVLM-navigation-dataset_finalInVLM-navigation-dataset_final_final2openapps-ui-tars-navigation-1500ep
OpenApps UI-TARS Navigation (1500 episodes)
UI-TARS-1.5-7B navigation trajectories on the OpenApps synthetic web-app suite, for world-model / GUI-agent research. Each episode is a short cross-app navigation: from a starting app, click Return to List of Apps to reach the launcher, then click the target app's tile. (A 150-episode subset is at taj-gillin/openapps-ui-tars-navigation-150ep.)
Schema (one row per step)
pixels — 1024x640 RGB screenshot the agent… See the full description on the dataset page: https://huggingface.co/datasets/taj-gillin/openapps-ui-tars-navigation-1500ep.251120_navigation_markopenapps-ui-tars-navigation-150ep
OpenApps UI-TARS Navigation (150 episodes)
UI-TARS-1.5-7B navigation trajectories on the OpenApps synthetic web-app suite, for world-model / GUI-agent research. Each episode is a short cross-app navigation: from a starting app, click Return to List of Apps to reach the launcher, then click the target app's tile.
Schema (one row per step)
pixels — 1024x640 RGB screenshot the agent observed.
action — discrete [type, grid_x, grid_y]: type 0=click / 1=scroll-down /… See the full description on the dataset page: https://huggingface.co/datasets/taj-gillin/openapps-ui-tars-navigation-150ep.navigationnavigation_aid
