CoolFace
Datasetpublic

Yangyihui/awm-nav-pretrain

AWM Navigation Pre-training Corpus (private) — uniform 30 fps Action-free video pre-training data for AWM. Egocentric go-to navigation clips, open-vocabulary route+landmark / object-level instructions (Gemini v3 harvest), across indoor (object-level) and outdoor (route+landmark) domains. All video is normalized to 30 fps (native 60fps down-sampled, native 30fps kept; 24/25fps clips dropped — no clean resample). fps_map.json records each clip's ORIGINAL fps for reference.… See the full description on the dataset page: https://huggingface.co/datasets/Yangyihui/awm-nav-pretrain.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
3likes4.9kdownloads
Dataset Card

AWM Navigation Pre-training Corpus (private) — uniform 30 fps

Action-free video pre-training data for AWM. Egocentric go-to navigation clips, open-vocabulary route+landmark / object-level instructions (Gemini v3 harvest), across indoor (object-level) and outdoor (route+landmark) domains. All video is normalized to 30 fps (native 60fps down-sampled, native 30fps kept; 24/25fps clips dropped — no clean resample). fps_map.json records each clip's ORIGINAL fps for reference.

videos/{indoor,outdoor}/<clip>.mp4     # 30 fps RGB clips
labels/{indoor,outdoor}/<clip>.json    # per-clip go-to segments [{t0,t_end,
                                        #   instruction, trajectory, ...}]
fps_map.json                           # clip -> original native fps

Video source: SpatialVID-HQ (FelixYuan/SpatialVID-HQ). Uploaded in parallel shards.

Yangyihui/awm-nav-pretrain · CoolFace