CoolFace
Datasetpublic

cy0307/awesome-egocentric-atlas

Use this dataset from datasets import load_dataset ds = load_dataset("cy0307/awesome-egocentric-atlas", split="train") print(len(ds), "resources") print(ds[0]) papers = load_dataset( "csv", data_files="https://huggingface.co/datasets/cy0307/awesome-egocentric-atlas/resolve/main/awesome-egocentric-papers.csv", split="train", ) print(len(papers), "paper-linked resources") Each row is one catalogued resource. Columns: Column Description name Resource name… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/awesome-egocentric-atlas.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
18likes2.9kdownloads
Dataset Card

Use this dataset

python
from datasets import load_dataset

ds = load_dataset("cy0307/awesome-egocentric-atlas", split="train")
print(len(ds), "resources")
print(ds[0])

papers = load_dataset(
    "csv",
    data_files="https://huggingface.co/datasets/cy0307/awesome-egocentric-atlas/resolve/main/awesome-egocentric-papers.csv",
    split="train",
)
print(len(papers), "paper-linked resources")

Each row is one catalogued resource. Columns:

ColumnDescription
nameResource name
kinddataset, benchmark, model, toolkit, or collection
releasedRelease date as YYYY or YYYY-MM
venuePublication venue (conference/journal or arXiv/project page)
statusAccess state: open, watch, partial, benchmark, or request
scopeegocentric (first-person) or adjacent (related, not first-person)
yearRelease year
urlPrimary link
paperPaper link, when available
codeCode link, when available
licenseLicense, when known
scaleOne-line description of scale / signal / contribution
tasksTask families (; -separated)
modalitiesModalities (; -separated)

The paper-focused sheet is also mirrored as `awesome-egocentric-papers.csv`, containing only resources with a paper link. A richer nested-JSON version with research-area summaries and stats is in `site-data.json`. This dataset card and catalog are mirrored from the GitHub repository and the interactive site; see the full README below for the resource tables and figures.


<p align="center"> <img src="assets/awesome-egocentric-atlas-cover.png" alt="Awesome Egocentric Atlas cover" width="100%"> </p>

<h1 align="center">Awesome Egocentric Atlas</h1>

<p align="center"> <img src="assets/awesome-egocentric-logo.png" alt="Awesome Egocentric Atlas logo" width="118"> </p>

<p align="center"> <strong>Datasets, benchmarks, models, and tools for egocentric AI.</strong> </p>

<!-- LANG-BAR:START --> <p align="center"> <a href="README.md"><b>English</b></a> · <a href="README.zh.md">中文</a> · <a href="README.es.md">Español</a> · <a href="README.fr.md">Français</a> · <a href="README.de.md">Deutsch</a> · <a href="README.ja.md">日本語</a> · <a href="README.ko.md">한국어</a> · <a href="README.pt.md">Português</a> · <a href="CONTRIBUTING.md">Help translate</a> · <a href="https://chaoyue0307.github.io/awesome-egocentric-atlas/">Landing page</a> · <a href="https://huggingface.co/datasets/cy0307/awesome-egocentric-atlas">Hugging Face mirror</a> </p> <!-- LANG-BAR:END -->

<p align="center"> <a href="https://github.com/sindresorhus/awesome"><img alt="awesome" src="https://awesome.re/badge-flat2.svg"></a> <a href="https://github.com/ChaoYue0307/awesome-egocentric-atlas/actions/workflows/validate.yml"><img alt="validate" src="https://github.com/ChaoYue0307/awesome-egocentric-atlas/actions/workflows/validate.yml/badge.svg"></a> <a href="https://chaoyue0307.github.io/awesome-egocentric-atlas/"><img alt="project site" src="https://img.shields.io/badge/site-GitHub%20Pages-067882"></a> <a href="https://huggingface.co/datasets/cy0307/awesome-egocentric-atlas"><img alt="Hugging Face mirror" src="https://img.shields.io/badge/Hugging%20Face-mirror-ffcc4d"></a> <a href="data/resources.yml"><img alt="resources" src="https://img.shields.io/badge/resources-866-0097A7"></a> <a href="README.md#dataset-atlas"><img alt="datasets" src="https://img.shields.io/badge/datasets-vision%20%7C%20robotics%20%7C%20memory-344054"></a> <a href="README.md#models-tools-and-baselines"><img alt="models and tools" src="https://img.shields.io/badge/models-and%20tools-F5A623"></a> <a href="LICENSE"><img alt="license" src="https://img.shields.io/badge/license-MIT-667085"></a> <a href="CONTRIBUTING.md"><img alt="PRs welcome" src="https://img.shields.io/badge/PRs-welcome-22A06B"></a> </p>

Awesome Egocentric Atlas is a practical catalog of egocentric (first-person) datasets, benchmarks, models, and tools for egocentric vision, embodied AI and robotics, vision-language-action, world models, long-context memory, AR/VR, and hand-object interaction. Every entry shows its public-access status, so you can tell at a glance what you can download today and what is still just a paper.

Updated: 2026-08-23. Scope: the main atlas is human or animal first-person capture from head, glasses, headset, body, wrist, handheld, or synchronized ego-exo rigs (where the ego view is central). Related but non-egocentric resources — robot-only datasets, multi-view robotic benchmarks, autonomous-driving 4D data, and general long-video reasoning — are listed separately under Adjacent and Related Resources rather than in the main tables.

<p align="center"> <img src="assets/awesome-egocentric-atlas-map.png" alt="Awesome Egocentric Atlas catalog overview" width="100%"> </p>

Contents

How to Use This Atlas

Use the first two tables to orient yourself, then jump into the detailed atlas section that matches your task. Each entry is deliberately short: name, release date (year-month, from the first public/arXiv posting where known), venue (conference/journal where published, or arXiv/source for preprints), scale/signal, best use, and public status. For filtering by task, modality, status, or date, use the catalog in `data/resources.yml`, where every entry carries a released field; label meanings and filter groups are explained in docs/taxonomy.md. Task-oriented research recipes walk you from goal to experiment.

Prefer a browsable view? The interactive site lets you filter the catalog by task, status, and date in eight languages, and the same catalog is mirrored as a dataset on Hugging Face.

<p align="center"> <img src="assets/awesome-egocentric-reader-route.png" alt="Awesome Egocentric Atlas reader route" width="100%"> </p>

At a Glance

SignalWhat it means for readers
866 egocentric resources229 datasets, 142 benchmarks, 442 models, and 43 toolkits, plus 4 collection hubs — across vision, robotics, memory, and AR. 178 related non-egocentric resources are listed separately.
6 research areasFoundation video, procedure/action, hands and 3D, memory/reasoning, robotics/VLA, and AR/wearable sensing.
5 access statesopen, request, benchmark, partial, and watch keep availability visible before you plan experiments.
Machine-checked catalog`data/resources.yml` is the source for type, year, status, URL, tasks, and provenance — and CI keeps the public artifacts in sync.
Reader-first tablesEach entry is short enough to scan, then links out to the official page, paper, code, or dataset portal.

<p align="center"> <img src="assets/awesome-egocentric-timeline.png" alt="Egocentric resources by era" width="100%"> </p>

Milestones

Representative works to read first, from early egocentric activity datasets to recent models and large-scale corpora.

<p align="center"> <img src="assets/awesome-egocentric-milestones.png" alt="Illustrated milestone timeline for representative egocentric AI works" width="100%"> </p>

<!-- MILESTONES:START (generated by build_artifacts.rb — do not edit by hand) -->

<!-- MILESTONES:END -->

Start Here

ReaderBest first moveThen go deeper
New to egocentric visionRead Ego4D, EPIC-KITCHENS-100, Ego-Exo4D, and EgoSchema first.Add EgoClip/EgoVLP for video-language, VISOR for masks, and HOT3D/HOI4D for 3D hand-object work.
Building a benchmark tableStart from the open and benchmark entries, then inspect request licenses.Use the YAML catalog to filter by task and modality.
Building an assistant or agentStart with Ego4D, Xperience-10M, EgoLife, TeleEgo, EgoSchema, MyEgo, EgoIntrospect, EgoBench, UCS-Bench, and Ego2Web.Track Ego-R1, X-LeBench, EgoMemReason, StreamMemBench, ReFocus, EgoCoT-Bench, EgoBlind, EgoSound, and NoRA for reasoning diagnostics.
Building a robot/VLA systemStart with Xperience-10M, EgoDex, EgoVerse, Ego-Exo4D, HoloAssist, HOI4D, HOT3D, Assembly101, OpenEgo, UMI, FastUMI-100K, EgoPAT3D, HUG, and WiYH.Add Dobb-E, FastUMI, MV-UMI, UMIGen, YUBI, Hoi!, EgoVLA, EgoEngine, HumanEgo, EgoGuide, Ego-Pi, EgoAERO, H2O, ARCTIC, HOMIE-toolkit, Xperience task baselines, and EgoHandTrajPred for policy/action supervision. For robot-only corpora see Adjacent and Related.
Studying hands, objects, and 3DStart with HOT3D, HOI4D, H2O, ARCTIC, FPHA, EgoHOS, EgoBody, GIMO, EgoHumans, EgoTactile, and UnrealEgo.Add EgoTracks/TREK-150 for tracking, EgoForce for monocular 3D hand pose, and EgoEMG/EgoEVHands/EgoPressDiff for emerging sensors and contact.

Landscape Snapshot

<p align="center"> <img src="assets/awesome-egocentric-task-matrix.png" alt="Egocentric AI task matrix" width="100%"> </p>

AreaWhat to compareCanonical anchors
Foundation videoscale, license, modalities, benchmark maturityXperience-10M, Ego4D, Ego-Exo4D, EPIC-KITCHENS-100, EgoClip
Procedure and actionaction granularity, mistakes, anticipation, gazeEPIC-KITCHENS, Assembly101, MECCANO, EgoProceL, HoloAssist
Hand-object and 3Dhands, objects, contact, pose, meshes, scene geometryHOT3D, HOI4D, H2O, ARCTIC, FPHA, EgoHOS
Memory and reasoningclip length, grounding, QA type, streaming constraintsEgoSchema, EgoLife, TeleEgo, MyEgo, X-LeBench, EgoMemReason, UCS-Bench, StreamMemBench, EgoSound, Ego-R1
Robotics and VLAaction labels, hand trajectories, robot transfer, task diversityXperience-10M, EgoDex, EgoVerse, OpenEgo, UMI, FastUMI-100K, YUBI, EgoPAT3D, HUG, WiYH, EgoVLA
AR/wearable sensinggaze, IMU, SLAM, audio, physiological signalsProject Aria, ADT, AEA, Nymeria, PROSE, EgoIntrospect, EasyCom
Origins. Egocentric vision grew out of wearable computing: Steve Mann pioneered continuous first-person ("sousveillance") cameras through the 1980s–90s and is widely credited with opening up the field, and Microsoft's SenseCam (2006) popularized lifelogging capture. The modern computer-vision wave began at the first IEEE Workshop on Egocentric Vision (CVPR 2009), where Spriggs, De La Torre, and Hebert introduced egocentric activity recognition on the CMU-MMAC database (2009) — the earliest dataset in this atlas — followed by the foundational GTEA (2011) and ADL (2012) datasets that launched first-person activity and object research.

Creator and Release Notes

ResourceCreated / stewarded byRelease channel and access note
Xperience-10MRopediaReleased under ropedia-ai on Hugging Face. Full dataset uses controlled non-commercial access; the sample is open, HOMIE-toolkit loads/visualizes annotations, and Xperience task baselines provide a public evaluation artifact.
Ego4DMeta AI / Facebook AI with an international academic consortiumOfficial data portal with license-gated downloads and benchmark tasks.
Ego-Exo4DMeta and academic partners across the Ego-Exo4D consortiumOfficial data portal with application/download tooling for synchronized ego-exo data.
EPIC-KITCHENSUniversity of Bristol-centered EPIC-KITCHENS teamOfficial challenge/dataset portal with public annotations and benchmark splits.
Project Aria datasetsMeta / Project AriaOfficial Project Aria dataset portal and tooling for AR-glasses sensor data.

Status Legend

StatusMeaning
openPublic download, public annotations, public code, or application-based access is clearly documented.
requestPublic project exists, but dataset access requires license, form, approval, or institutional agreement.
benchmarkMainly a benchmark, challenge, labels, or evaluation over existing videos.
partialSome assets are public, but full raw data, license, or long-term availability is unclear.
watchRecent or important resource whose release state still needs verification.

For how each status is decided, see docs/status_policy.md.

<p align="center"> <img src="assets/awesome-egocentric-access-funnel.png" alt="Egocentric resources by access state" width="100%"> </p>

Fast Task Map

Research goalStrong starting points
General egocentric foundation dataXperience-10M, Ego4D, Ego-Exo4D, EPIC-KITCHENS-100, EgoClip
Ego-exo and skill learningEgo-Exo4D, EgoExoLearn, Assembly101, Look and Tell
Robot manipulation / VLAXperience-10M, EgoDex, EgoVerse, OpenEgo, UMI, FastUMI-100K, YUBI, HUG, EgoVLA, EgoEngine, HumanEgo
Long-context QA and memoryEgoSchema, EgoLife, EgoMemReason, UCS-Bench, StreamMemBench, Ego-EXTRA, EgoStream, EgoBench
Hand-object and 3D HOIHOT3D, HOI4D, EgoHOS, FPHA, HUG, EgoTactile, EgoForce, EgoEMG
Audio, speech, and social interactionEgo4D AV tasks, EasyCom, EgoCom / audio-visual correspondence, AV-CONV
AR glasses and scene understandingProject Aria Datasets, Aria Digital Twin, AEA, Nymeria, DTC

Recipes and Reference

Task-oriented starting points and the references that keep the catalog usable.

Research recipes — each gives a goal, what to start with, what to add, benchmarks, baselines, a minimum viable experiment, common pitfalls, and a reporting checklist:

  • —Long-context egocentric video QA
  • —Robot / VLA from egocentric human video
  • —Hand-object 3D HOI
  • —AR-glasses and wearable sensing
  • —Audio and social egocentric understanding

Reference

  • —Taxonomy — task/modality families and how to filter the catalog.
  • —Access Matrix — data, code, license, and leaderboard per key resource.
  • —Status Policy — how open/request/benchmark/partial/watch are decided.
  • —Resource Schema — every field in data/resources.yml.
  • —Maintenance — the add → verify → validate workflow.

Dataset Atlas

Foundational and Large-Scale Ego Video

Broad, large-scale first-person corpora that most egocentric pipelines pretrain or benchmark on.

ResourceReleasedVenueScale / signalBest forStatus
RekaDaily-10k (raw)2026-08Hugging FaceCurrent incremental release: 7,834 hours, 397,171 unscripted first-person videos, 9,836 WebDataset shards, and about 70 TBLarge-scale daily-life pretraining, action understanding, and human-to-robot learningopen
Ego5002026-08Hugging FaceAnnounced 500-hour corpus with 310K+ structured action annotations across 390 task-environment combinations and 60+ environments; manual access and upload are still pendingDense work-procedure, hand-state, and task-phase modelingrequest
EgoSuite-Open100K2026-08Hugging FaceGated EgoDemo, EgoStandard, and EgoPro buckets advertise 100K hours across head and head-plus-wrist lines with hand/body pose and semantics; 100K remains the collection target, while current manifests are the release truthLarge-scale controlled-access egocentric pretraining and VLA datarequest
Stereo-5502026-07Hugging Face550 hours per camera, 1,462 calibrated stereo sessions, 209,315 dense action segments, and synchronized 6-axis IMU for 1,271 sessionsCalibrated stereo HOI, dense video-language learning, and sensor fusionopen
ACE-Data-02026-07arXivCurrent public snapshot has 30 room-scale takes, 95,238 files, and about 3.61 TB of ego/exo video and multimodal sidecars; 150 hours and 75K+ interactions remain the full-release targetMultisensory household imitation, world models, and VLA researchpartial
RetailSMV2026-06arXiv32,105 captioned synchronized retail clips from five supermarkets with staff-perspective egocentric and exocentric viewsRetail world-model adaptation and ego/exo viewpoint comparisonwatch
EgoCS-400K2026-06arXiv400K+ first-person Counter-Strike gameplay videos (10K hours, 13 maps) with aligned actions, player state, camera motion, and game eventsAction-conditioned interactive world models from first-person gameplaywatch
HumanNet2026-05arXiv1M-hour human-centric video corpus spanning first-person and third-person views, with captions, motion descriptions, hand/body signals, and a reported 1K-hour egocentric subset for VLA-style validationLarge-scale human-video pretraining, interaction understanding, and human-to-robot transferwatch
EgoInteract2026-05arXivPublic synthetic release with 10,534 episodes, about 1.9M frames, 61,211 NAO/HOI task frames, and released simulator/training codeSynthetic temporal segmentation, next-active-object detection, anticipation, and EHOIopen
PRISM2026-03arXivPaper reports 270K multi-view retail SFT samples; the gated PRISM-100K release provides 100K samples, 25,173 ego/exo videos, 21 task types, and about 772 GBEmbodied VLM spatial, physical, and action reasoning in retail settingsrequest
Ego-1K2026-03CVPR 2026Nearly 1,000 synchronized multiview egocentric videos from a custom 12-camera plus VR-headset rigDynamic 3D/4D scene understanding and novel view synthesis from ego rigsopen
EgoCrowds / CrowdEraser2026-03arXivSemi-synthetic paired crowded/empty clips from real egocentric walking-tour video; CrowdEraser diffusion removes crowds for humanless walkthroughsFirst-person walking-tour video editing and environment modelingwatch
Xperience-10M2026-03Hugging FaceRopedia release on Hugging Face; 10M experiences, 10K hours, six video streams, audio, stereo depth, camera pose, hand/body mocap, IMU, hierarchical language, ~1 PB totalEmbodied AI, world models, robot learning from human experience, sensor fusion, 3D/4D understandingrequest
Xperience-10M Sample2026-03Hugging FacePublic sample episode for Xperience-10M with six rows on Hugging Face and cc-by-nc-4.0 termsLoader testing, demos, task-suite prototyping, annotation inspectionopen
Egocentric-100K2025-12Hugging Face100,405-hour manual-labor egocentric video corpus with 10.8B frames, 2.01M clips, intrinsics, and WebDataset shardsIndustrial/manual-labor pretraining, robot learning, and active manipulation density analysisrequest
Egocentric-10K2025-11Hugging Face10,000-hour factory-only egocentric video corpus with 1.08B frames, intrinsics, and WebDataset-style shardsIndustrial egocentric pretraining and active-manipulation density analysisrequest
Voxel51 Egocentric-10K Subset2025-11Hugging FaceOpen FiftyOne package for a Factory 51 subset of Egocentric-10K, with the card reporting 416 1080p head-mounted industrial clipsLightweight inspection, prototyping, and visualization of Egocentric-10K-style factory footageopen
Look and Tell2025-10NeurIPS 2025 Workshop25 participants, Project Aria plus stationary cameras, gaze/speech/video, 3D reconstructionsReferential communication across ego and exo viewpointswatch
EgoBlind2025-03NeurIPS 2025 D&B1,392 first-person videos from blind and visually impaired users, 5,311 questions posed or verified by blind usersAssistive egocentric VideoQA for blind userswatch
HD-EPIC2025-02CVPR 202541 hours, 9 kitchens, dense fine-grained labels, 3D fixture annotations, audio events, VQADetailed kitchen understanding, VQA, 3D-aware ego reasoning, audio-event recognitionopen
EgoVid-5M2024-11NeurIPS 20255M curated egocentric clips at 1080p with fine-grained kinematic and high-level text action annotations (NeurIPS 2025)Egocentric video generation and action-conditioned world modelsopen
Ego-Exo4D2023-11CVPR 20241,286 hours, 740 participants, synchronized first- and third-person views, audio, gaze, IMU, 3D point clouds, language commentarySkilled activity, ego-exo transfer, cross-view translation, pose, proficiency, roboticsrequest
HoloAssist2023-09ICCV 2023166 hours, 350 instructor-performer pairs, 7 synchronized streams from mixed-reality headsetsInteractive assistance, mistake detection, intervention prediction, procedural collaborationopen
EgoObjects2023-09ICCV 2023Pilot release with 9K+ videos, 250 participants, 50+ countries, 650K annotations, 368 categoriesCategory and instance-level egocentric object understanding, continual object detectionopen
EgoCom / Ego audio-visual correspondence2023-07CVPR 2024Egocentric video with spatial audio for conversation and audio-visual correspondence tasksActive speaker detection, spatial audio denoising, conversational graph reasoningopen
Assembly1012022-03CVPR 20224,321 videos, 8 static plus 4 egocentric views, 100K coarse and 1M fine action segments, 18M 3D hand posesProcedural action, anticipation, mistake detection, cross-view transferopen
Ego4D2021-10CVPR 20223,670+ hours from 900+ camera wearers, multiple countries, video/audio/gaze/stereo/3D/narrations depending on subsetLong-form ego video, episodic memory, social, hand-object, forecasting, audio-visual tasksrequest
EasyCom2021-07arXivAR glasses egocentric multi-channel audio and wide-FOV RGB for noisy conversationsSpeech enhancement, source localization, conversation assistanceopen
EPIC-KITCHENS-1002020-06IJCV 2022100 hours, 20M frames, 90K actions, 45 kitchens, narrations and dense action labelsKitchen action recognition, anticipation, action detection, retrievalopen
From Third Person to First Person2018-12CVPR 2019Ego/exo video synthesis and retrieval datasets plus baselines for bridging third-person social video to first-person viewsCross-view synthesis, retrieval, and ego-exo transferwatch
EPIC-KITCHENS2018-04ECCV 201855 hours / 11.5M frames from 32 kitchens with narrations, 39.6K action segments, and object boxesOriginal large-scale kitchen action recognition and anticipation benchmarkopen

Robotics, Manipulation, and VLA

Egocentric human-manipulation and wrist-camera data aimed at vision-language-action models and human-to-robot transfer.

ResourceReleasedVenueScale / signalBest forStatus
Egocentric Manufacturing2026-08Hugging Face259 gated clips, about 100 hours, 14 live-factory process domains, 256 workers, 3,911 scenes, and 2,081 entitiesIndustrial procedures, VLA pretraining, and authentic plant-floor HOIrequest
Axis Ego Samples2026-08Hugging FacePublic 100-hour LeRobot egocentric sample across nine-plus work domains, plus ten episodes from a separate 150-hour delivery corpus; reuse terms are unstatedBroad workplace human-demonstration and loader experimentsopen
RealFactory-Ego-Data2026-08Hugging FaceFive Apache-2.0 head-mounted episodes totaling about 14 minutes on a real touchscreen line, with RGB, 21-point hands, and frame-level actionsCompact industrial HOI and action-segmentation prototypesopen
BaseMatrix EGO Desktop Demo2026-08Hugging FaceOne 70-second, 2,114-frame binocular ego/exo episode with camera pose, metric 3D hands, contact, and derived actionsInspecting a dense stereo hand-retargeting schemaopen
ActTrace Ego-Home2026-08Hugging Face13 household episodes over 11 tasks and 18 minutes with RGB, LiDAR depth, 100 Hz IMU, metric hands, and IMU-verified 6-DoF head trajectoriesMetric RGB-D/IMU household demonstrations and sensor fusionopen
Everyday Manipulation 3D2026-08Hugging Face116 densely annotated episodes and 28,019 frames with depth, pose, MANO, contact, masks, and language; companion raw release has 1,513 clips, 10.28 hours, and 279 GiBDense 3D HOI evaluation plus raw RGB-D pretrainingopen
VITRA Egocentric Instruction Sidecars2026-08Hugging Face95,831 accepted hand-action instructions from 175,327 Qwen3.5 candidates across GigaHands, HOT3D, and OakInk2; source revisions and paths are preserved, but no standalone license is declaredVLA instruction tuning and annotation-sidecar researchpartial
InHandPlus Wrist-View Camera and IMU Sample2026-08Hugging FaceAuto-gated 30-episode cloth-manipulation sample with wrist video, synchronized raw 6-axis IMU, and LeRobot v3 metadata under CC-BY-NC-4.0Wrist-view sensor fusion and early loader experimentsrequest
InfoBayAI Egocentric Video Sample2026-08Hugging Face10 gated MP4 samples and metadata totaling 31.87 GB; the much larger headline collection on the card is not present in the repository manifestInspecting first-person sample media without treating the advertised corpus as releasedrequest
UniDataPro Egocentric Video Sample2026-08Hugging FaceOne public 2.33 GB video/metadata/tracking preview for a commercial collection advertised as 4,050 hours of Pico, ZED, IMU, and motion-tracker captureEvaluating a commercial multimodal manipulation-data offeringrequest
Nexdata 10,000-Hour Egocentric Video Dataset2026-08Hugging FaceCommercial specification for 10,000 hours of 4K stereo head video, wrist/ankle IMU, 76-point body pose, and dense actions; full media is available by requestEvaluating a large commercial multimodal capture offeringrequest
ROCO IROS 2026 UMI Dataset2026-08Hugging Face4,928 public LeRobot episodes and 7.46M frames with stereo FPV, four hand cameras, and bimanual pose/gripper state; no card or license is suppliedBimanual UMI/VLA loading and six-camera policy experimentspartial
DreamTraj / MOVE2026-08arXiv5,038 language-conditioned egocentric 6-DoF object trajectories; predicts motion from unrendered video-diffusion latents at 4.6x the reported speed of generate-then-extract pipelinesObject-trajectory prediction and language-to-manipulation motionwatch
Ego2Robot2026-08arXivConverts egocentric human manipulation video into 18,561 hours of synthesized robot-format data across 15 morphologiesScalable ego-to-robot pretraining and embodiment generalizationwatch
RoboReact2026-08arXivGenerates egocentric RGB-D manipulation plans, reconstructs interaction keyframes, retargets them to humanoids, and refines execution in closed loopGenerated-video-to-humanoid whole-body skill distillationwatch
SiMDex2026-08arXivMines 1.49M task-relevant samples from roughly 32M egocentric clips and reports 47.7% to 61.1% dexterous success over equal-sized random samplingRelevance-aware human-video curation for dexterous VLA post-trainingwatch
VLAff / EgoAffordance2026-08IROS 2026204K episodes with 5.6M visual and 11.6M grasp/trajectory affordances; predicts heatmaps, grasps, and executable trajectoriesActionable affordance learning and human-to-robot transferwatch
JoyAI-RA 0.52026-08arXivDual latent/canonical action alignment across human egocentric video, simulation, and robot data for a generalist VLWA policyScaling manipulation learning from heterogeneous human and robot experiencewatch
HumynLabs Egocentric Sample Collection2026-08Hugging FaceFive CC-BY-4.0 packs covering head/dual-wrist video and IMU, stereo/mono labels, voice-over captions, and regional residential IMUSmall multimodal capture, sensor-fusion, and loader experimentsopen
EgoGenesis2026-07arXivGeometry-aware egocentric world-action simulator; 400 generated plus 400 real trajectories improve reported OOD success from 77% to 84% single-arm and 53% to 70% dual-armControllable ego-video generation and world-action-model augmentationwatch
Firstly Household Manipulation v0.12026-07Hugging Face12 open dishwashing/laundry episodes totaling 119.1 minutes and 214,437 frames, with 148 phase instructions and two-hand tracksCompact LeRobot human-demonstration and annotation-pipeline testingopen
HiFi-UMI-2K2026-07arXiv2,000 open hours from a high-fidelity UMI rig with head stereo-inertial SLAM, native inter-gripper pose, microsecond synchronization, and two ultra-wide cameras per handRobot-free VLA/WAM pretraining and direct manipulation-policy deploymentopen
EgoRecovery2026-07arXivEgocentric human recovery demonstrations collected at more than 10x robot-teleoperation throughput and aligned with a shared corrective-intent spaceLearning when and how robots should recover from manipulation failureswatch
DDD Egocentric Stereo Manipulation Sample2026-07Hugging FaceTen open LeRobot v2 episodes (about 2.37 GB) with stereo first-person video and high-frequency hand/head poseStereo human-demonstration loaders, hand tracking, and sensor fusionopen
Showway Egocentric Origami Series2026-07Hugging FaceNine open neck-mounted iPhone origami episodes across three LeRobot v3 repositories, with 3D hand joints and reviewed narrationCompact procedural HOI and imitation-learning experimentsopen
EGXO Household Egocentric Video Evaluation2026-07Hugging Face71 household-task videos totaling 10 hours; six previews are public and the complete 43.4 GiB media package is licensed by requestHousehold procedure and hand-object video evaluationrequest
Let the Body Follow2026-07arXivCoupled egocentric TIAGo teleoperation maps head and arm motion to coordinated torso and mobile-base control, reducing explicit controls and user workloadWhole-body mobile-manipulator teleoperation and HRIwatch
EgoHTR2026-07arXiv55 scene-aligned 4D human-terrain sequences across seven environments, totaling 1.37 hours and about 150K frames from eight participants, with multiview and mocap ground truthScene-aware human-motion reconstruction and humanoid terrain traversalwatch
Open-AoE2026-07Hugging FaceNano 3-hour and tiny 100-hour subsets are released; about 323 hours were uploaded and 694 hours prepared by August 12, with 11.9 TB currently stored; 2,000 hours remains the roadmapLarge-scale hand-motion, camera-trajectory, and manipulation pretrainingpartial
SenseXperience Raw MCAP Sample2026-07Hugging Face12 public raw ROS 2 MCAP episodes with four synchronized cameras, head IMU near 120 Hz, calibration, metadata, and recorder diagnosticsRaw multimodal capture-pipeline and sensor-synchronization testingopen
Ego Data by Object2026-07Hugging Face4,110 pick-and-place episodes over 51 object labels with egocentric RGB and HDF5 hand/body transforms, intrinsics, task attributes, and descriptions; source licensing is unclearObject-conditioned hand tracking and imitation-data experimentspartial
EgoExo-Wreck2026-07Hugging FaceCommercial card reports 150+ synchronized ego-exo hours of object destruction, debris handling, cleanup, and reset; only a gated 1 GB evaluation sample is hostedPhysics-rich HOI, state-change, world-model, and cleanup demonstrationsrequest
Verbose Industrial Egocentric Sample Series2026-07Hugging FaceSix open industrial sample repositories with 53 first-person clips, 4.86 hours, and about 17.0 GB across textile, metal, electronics assembly, skilled work, and manufacturing operationsCompact industrial action and human-demonstration experimentsopen
Verbose Cooking and Chopping Egocentric Video2026-07Hugging Face11 head- or chest-mounted kitchen clips totaling 101.7 minutes in a public 6.09 GB ZIP, with audio and clip metadata across cooking and choppingOpen kitchen action, procedure, and hand-object video experimentsopen
Verbose Household Cleaning Egocentric Video2026-07Hugging Face12 head- or chest-mounted 1080p/30 fps cleaning clips totaling 104.3 minutes in a public 5.74 GB ZIP, with audio and three sub-activity labelsOpen domestic-action and assistive-robotics video experimentsopen
Flikforge Egocentric Household Tasks2026-07Hugging FaceCard claims about 100 rights-cleared household clips with dense customizable annotations, but the repository currently has no media or annotation files and its LICENSE is emptyMonitoring a prospective commercial household-demonstration sample releasewatch
AgenticFocus2026-07arXivRestores occluded object geometry and full-hand motion from ordinary first-person human videos, then retargets and composites robot-aligned observations, actions, and statesMixed-reality human-to-humanoid demonstration synthesiswatch
Hub Egocentric GoPro RGB+IMU2026-07Hugging Face19 GoPro HERO13 manipulation clips totaling 3.72 hours, with RGB, approximately 200 Hz GPMF IMU, sensor sidecars, and MCAPHigh-rate visual-inertial human-demonstration prototypingpartial
Hub Egocentric iPhone RGB+IMU2026-07Hugging Face12 iPhone manipulation clips totaling 2.96 hours, with 1080p RGB, CoreMotion/ARKit IMU and attitude, intrinsics, sidecars, and MCAPSmartphone RGB-IMU synchronization and sensor-fusion experimentspartial
Hub Egocentric Stereo RGB-D2026-07Hugging Face9 ZED X Mini clips with stereo RGB, metric depth, IMU, 6DoF VIO, calibration, MCAP, and per-clip LeRobot v3.0 packagesRGB-D/VIO pipelines and multimodal imitation-learning loaderspartial
Hub Egocentric iPhone RGB2026-07Hugging Face40 household-manipulation clips totaling 1.60 hours, delivered as 1080p H.264 RGB without IMU, depth, pose, or robot actionsLightweight first-person manipulation-video baselinespartial
Datoric Residential Egocentric Video2026-07Hugging FaceCommercial card claims 100,000 hours of 1080p+ household first-person video from Brazilian homes with task, action, object, HOI, and completion-state annotations; public files are specification/sample metadataReviewing large-scale household capture specifications and requesting production datarequest
Datoric Industrial Egocentric Video2026-07Hugging FaceCommercial card claims 50,000 hours of workplace first-person video across logistics, inspection, assembly, and cleaning; public files are specification/sample metadataReviewing industrial capture specifications and requesting production datarequest
EgoWAM2026-07arXivControlled human-robot WAM co-training study comparing pixel, DINO, and 3D-flow targets on three real bimanual tasks, with up to 4x OOD gains reported for DINOHuman-to-robot world-representation transferwatch
USA Egocentric2026-07Hugging FaceThree iPhone kitchen-manipulation episodes with RGB videos, camera trajectories, hand outputs, action segments, object/contact tracks, and quality reportsOpen manipulation-pipeline debugging and VLA annotation prototypingopen
RynnWorld-4D2026-07arXiv4D embodied world model over future RGB, depth, and optical flow; reports Rynn4DDataset 1.0 with 254.4M+ frames across egocentric human and robotic manipulation videosRobot manipulation world modeling and action-conditioned 4D predictionwatch
RynnWorld-Teleop2026-07arXivDigital-teleoperation world model driven by operator hand-pose streams to synthesize egocentric videos and action labelsScalable bimanual robot data generation from human pose streamswatch
LingBot-VLA 2.02026-07arXivVLA update trained with around 60K hours, including 10K hours of egocentric human videos and 50K hours of robot trajectoriesWhole-body VLA pretraining and practical robot deploymentwatch
LingBot-Video2026-07arXivMoE DiT video foundation model for embodied intelligence, adding robot-oriented manipulation/navigation and egocentric-perspective footage plus physically grounded rewardsEmbodied video pretraining and robot/world-model generationwatch
LIME2026-07arXivIntent-aware camera-motion generator mined from egocentric video, pairing language intents with relative SE(3) target posesActive camera control and next-view prediction for embodied agentswatch
H-Tac / TTP2026-07arXiv160-hour egocentric human tactile-action dataset with 300+ tasks and 135K episodes, plus transferable tactile pretrainingContact-rich dexterous manipulation from human tactile demonstrationswatch
HomER v22026-06Hugging Face765 public first-person household videos totaling 99.60 hours from 420 participants across 49 countries and 10 activity categoriesDiverse everyday activity, hand-object interaction, video-language, and robot-learning researchopen
Ego-Tactile Manipulation2026-06Hugging Face72 public episodes totaling 1.28 hours and 138,629 frames with 1080p ego video, 160-channel dual-hand pressure, head/wrist IMU, and contact-derived action segmentsOpen tactile-grounded manipulation, contact, and action-segmentation researchopen
EgoSteer2026-06arXivOpen full-stack dexterous VLA system using 9.6K hours of curated egocentric human video, with EgoSmith, a robot stack, and two released 3B checkpointsSteerable dexterous manipulation and human-video-to-robot pretrainingopen
Poseidon First-Person Task Video Samples2026-06Hugging FaceFive auto-gated 1080p clips totaling 174 seconds across kitchen, gardening, maintenance, fine-motor, and cleaning tasksReviewing task diversity, hand-object visibility, and capture qualityrequest
WT-UMI2026-06arXivWearable whole-body tactile UMI interface aligning human demonstrations and teleoperation through tactile images, contact forces, and end-effector posesContact-aware humanoid whole-body manipulationwatch
Human-as-Humanoid2026-06arXivSynchronized ego-exo human videos converted through motion recovery, IK, and FK-aware supervision into 60-DoF humanoid action chunksZero-shot humanoid VLA learning from human demonstrationswatch
UMI-Bench 1.02026-06arXivLocal-first real-robot evaluation protocol dedicated to UMI-style data collection, deployment, reset, logging, and task-factor analysisReproducible UMI policy benchmarkingwatch
VISTA UMI2026-06arXivAdapts UMI data for VLA training with wrist-fisheye UMI-VQA visual grounding plus physics-based trajectory validationMaking UMI demonstrations usable for VLAswatch
Pose6DAug2026-06arXivPhysically plausible multi-view object-swapping augmentation for VLA policies under occlusion and egocentric viewpointsRobust manipulation data augmentationwatch
MotionWAM2026-06arXivReal-time humanoid world-action model driven by a single egocentric camera and unified whole-body motion tokensHumanoid loco-manipulation from egocentric perceptionwatch
OSCAR2026-06arXivOmni-embodiment action-conditioned video world model trained across robotics and egocentric human datasetsRobot policy evaluation and cross-embodiment world modelingwatch
Mem-World2026-06arXivWrist-view-centered 4D surfel memory for persistent action-conditioned manipulation world modelsLong-horizon wrist-camera world modelingwatch
PAIWorld2026-06arXiv3D-consistent multi-view world foundation model across egocentric, eye-to-hand, and wrist camerasMulti-view robotic world modelingwatch
Qwen-RobotManip2026-06arXiv38,100-hour manipulation pretraining corpus converting egocentric hand demonstrations into robot trajectories across 15 platformsScalable VLA robot manipulationwatch
CAIP2026-06arXivContrastive action-image pretraining from 32,041 hours of egocentric human video with 3D hand-keypoint action proxiesAction-centric visual pretraining for roboticswatch
YUBI2026-06arXivReported 8,434 hours, 1.20M episodes, and 119 tasks from a finger-aligned UMI-style bimanual interfaceScalable bimanual dexterous data collectionwatch
HumanoidUMI2026-06arXivVR-assisted egocentric data capture for whole-body humanoid imitation, using sparse human keypoint trajectories and wrist-view observations for whole-body transferWhole-body humanoid transfer from egocentric human demonstrationswatch
BRIDGE State-Gated Experts2026-06arXivCombines handheld UMI demonstrations with targeted teleoperated segments through state-gated diffusion-policy expertsContact-rich manipulation with mixed UMI and teleoperation supervisionwatch
ForceBand2026-06arXiv10-hour wrist-worn sEMG, egocentric video, IMU, and fingertip-force dataset for force-enriched demonstrationsForce-aware human demonstration collection for robot policieswatch
EgoEngine2026-06arXivConverts egocentric human manipulation videos into high-fidelity robot observation videos and executable robot action trajectoriesHuman-to-robot data generation, dexterous imitationwatch
EgoAERO2026-06arXivAsset-free conversion from a single egocentric RGB-D demonstration; introduces EgoDex-R in the paperSingle-demo dexterous robot learningwatch
1M-HUGs / HUG2026-06arXiv1M frames / 27.8 hours of smart-glasses human grasps over 6,707 object instances, plus HUG-Bench with 90 unseen objectsHuman grasp modeling and zero-shot robot graspingopen
HALOMI2026-06arXivExtends UMI-style demonstration collection with egocentric head/wrist observations and head-hand trajectories for humanoid loco-manipulationActive-perception humanoid manipulation from human demonstrationswatch
HumanoidArena2026-06arXivBenchmark for egocentric hierarchical whole-body learning across seven leg-critical humanoid-object and humanoid-scene interaction tasksWhole-body humanoid control from egocentric perceptionwatch
EgoStation Smartphone Raw Catalog v12026-05Hugging FaceMetadata for 5,962 episodes and 693.9 hours, with 16 downloadable full-quality samples carrying raw video, hand keypoints, and depth overlaysInspecting smartphone capture quality and requesting selected raw demonstrationspartial
Luel IMU Egocentric 250k2026-05Hugging FaceCommercial card reports 250K hours and 46,340 tasks; the gated Hub repository exposes a verifiable 10-clip, 30-minute sample plus the full task taxonomyReviewing industrial RGB-IMU capture scope and requesting the proprietary full corpusrequest
RoboX-EgoTask2026-05Hugging Face17 public RGB-D task clips from five recordings with hand signals, camera pose, IMU, action segments, trajectories, segmentation, and quality reportsOpen multimodal hand-object interaction and human-demonstration prototypingopen
SenseXperience UMI Human Demonstrations2026-05Hugging Face1,051 public human-demonstration episodes, 220,963 frames, and 92 tasks with six synchronized head, wrist, depth, and gripper views plus 7-DoF poses in LeRobot formatOpen human-to-robot transfer, bimanual manipulation, and VLA dataopen
tau0-WM2026-05arXivUnified video-action world model trained on real-robot teleoperation, UMI-style interaction, egocentric human video, and rollout/failure dataFuture-aware robot action generation and action-conditioned simulationwatch
SABER2026-05arXiv100+ hours of in-store retail egocentric action capture paired with 360-degree exocentric contextRetail-domain VLA adaptation and manipulation data scalingwatch
Qwen-VLA2026-05arXivUnified VLA model trained over manipulation trajectories, egocentric human demonstrations, simulation, navigation, and trajectory supervisionGeneral embodied action modelingwatch
Gaze2Act2026-05arXivMaps first-person gaze into robot perspective for target specification in interactive manipulationGaze-conditioned VLA manipulationwatch
DeMiAn2026-05arXivDense multi-aspect language re-annotation over 1M robot clips and 50K EgoVerse human-egocentric videosLanguage-dense robot policy learningwatch
Grid Egocentric Residential Samples2026-05Hugging FacePublic first-person residential manipulation preview pack with full-length 1080p household task videosHousehold VLA capture-quality review and imitation-learning samplesopen
Industrial Workplace Egocentric FHD Samples2026-05Hugging Face21 rights-cleared head-mounted workplace clips with paired JSON metadata across factory, construction, services, retail, transit, and food-service domainsIndustrial and vocational human-demonstration samples for VLA/WAM prototypingopen
UniDataPro Egocentric Video Preview2026-05Hugging FaceCommercial preview card for a claimed 4,050-hour VR-headset/Zed first-person robotics collection; public files expose a sample video, CSV, and tracking textMarket scan and modality reference for large-scale egocentric robotics capturepartial
Lo6yu Egocentric RGB-D + EMG/IMU2026-05Hugging FaceEight household activity packages with RGB-D, wrist EMG/IMU, hand keypoints, masks, contact, per-finger force, semantic segments, and training-ready exportsMultimodal contact-aware imitation learning and sensor-fusion robot dataopen
Egocentric Adjust Bottle LeRobot2026-05Hugging FaceLeRobot-format human_mano adjust-bottle task with 500 episodes and 36,897 frames across image, text, timeseries, and video exportsLoader tests and small-scale LeRobot manipulation policy prototypingopen
Low-Resolution Active Perception BC2026-05arXivBehavior cloning with low-resolution wrist-mounted egocentric RGB for object finding and grasp triggeringActive perception under low computewatch
EgoSPT / SPOT2026-05arXivEgocentric spatially prompted manipulation trajectories with first-frame object/target grounding and 3D end-effector motionSpatially grounded manipulation trajectory predictionwatch
Mobile UMI2026-05arXivRobot-free mobile manipulation capture with chest-centric context, wrist-centric interaction, decoupled kinematics, and latency-aware diffusion policyMobile manipulation from portable demonstrationswatch
BifrostUMI2026-05arXivVR-device robot-free demonstrations with sparse keypoint trajectories, wrist-mounted visual data, and humanoid retargetingHumanoid whole-body demonstration collectionwatch
UniT2026-04arXivUnified latent action tokenizer anchoring massive egocentric human data to humanoid policy learning and world modelingHuman-to-humanoid transferwatch
XRZero-G02026-04arXivVR and dual-gripper robot-free collection system with a reported 2,000-hour dataset and data-mixing studyScalable robot-free manipulation data collectionwatch
EgoVerse2026-04arXiv1,362 hours, 80K episodes, 1,965 tasks, 240 scenes, 2,087 demonstratorsHuman demonstration scaling for robot learning and VLAwatch
RoboX Egocentric Collection2026-04Hugging Face7,342 auto-gated first-person clips across grasping, daily activities, scene capture, and navigation with metadata, hand keypoints, object tracks, action segments, IMU, and camera poseMulti-task robotics imitation data and navigation/manipulation benchmark prototypingrequest
EgoLive2026-04arXivLarge-scale real-world task-oriented egocentric routines for robot manipulationHome service, retail, and real-world work-task manipulationwatch
GazeVLA2026-04arXivVLA policy that pretrains on large-scale egocentric human data to capture gaze, intention, and action before robot fine-tuningGaze- and intent-aware robot manipulationwatch
WARPED2026-04arXivWrist-aligned rendering converts monocular egocentric human demonstrations into robot policy observations with 3D Gaussian SplattingCross-embodiment imitation from human ego videowatch
UMI-3D2026-04arXivExtends UMI with wrist LiDAR, LiDAR-centric SLAM, synchronized sensing, and spatiotemporal calibration for robust 3D demonstration capture3D-aware UMI data collectionwatch
OmniUMI2026-04arXivAdds RGB, depth, trajectory, tactile sensing, grasp force, and external wrench signals to a human-aligned UMI-style handheld systemContact-rich multimodal UMI demonstrationswatch
HRDexDB2026-04arXiv1.4K human/robot grasping trials, tactile, multiview video, egocentric video streamsCross-domain dexterous grasp learningwatch
UMI-Underwater2026-03arXivTransfers on-land handheld human demonstrations to underwater grasping through depth-based affordance representationsUnderwater manipulation without underwater teleoperationwatch
HoMMI2026-03arXivWhole-body mobile manipulation interface augmenting UMI with egocentric sensing, relaxed head actions, and cross-embodiment hand-eye policy designRobot-free mobile manipulation demonstrationswatch
Ego-Exo Manufacturing2026-03Hugging FaceGated 275-hour shoe-manufacturing corpus with 40 synchronized ego-exo groups, 80 standalone ego sessions, 137 exo sessions, and procedure/proficiency/mistake annotationsIndustrial ego-exo learning, skill assessment, and procedural modelingrequest
OBayData Egocentric Dexterous Manipulation Demo2026-03Hugging Face500 train and 100 test first-person manipulation sessions with head and wrist videos, 3D hand joints, camera extrinsics, action labels, and language annotationsQuality review and prototyping for dexterous human-demonstration pipelinesopen
HuMI2026-02arXivPortable robot-free whole-body demonstration interface with open code and seven Hugging Face datasets across kneeling, squatting, tossing, walking, and bimanual tasksHumanoid whole-body manipulation from portable human demonstrationsopen
Seesaw Video Subset2026-02Hugging Face763 gated LeRobot-compatible indoor episodes, including 243 ego and 520 exo views, 321,178 frames, four room settings, and smartphone AR camera poseEgo-exo robot learning, action recognition, and data-conversion experimentsrequest
CoMe-VLA2026-02arXivCognitive and memory-aware VLA that learns active-perception strategies from large-scale egocentric human dataNon-Markovian active perception and manipulationwatch
EgoAVFlow2026-02arXivLearns manipulation and active camera control from egocentric human videos through shared 3D flowActive-vision robot policy transferwatch
EgoHumanoid2026-02arXivCo-trains humanoid VLA policies from robot-free egocentric human demonstrations and limited robot dataHumanoid loco-manipulationwatch
EgoActor2026-02arXivGrounds high-level instructions into spatially aware egocentric humanoid actionsHumanoid task planning and action groundingwatch
AoE: Always-on Egocentric2026-02arXivAlways-on egocentric human-video collection pipeline and corpus for embodied AIScaling human-video data for robot learningwatch
EgoScale2026-02arXiv20,854 hours of action-labeled egocentric human video with a human-to-robot two-stage transfer recipe and a log-linear data-scaling lawScaling dexterous manipulation from human videowatch
Manus Egocentric Sample2026-01Hugging Face20 LeRobot-format manipulation episodes with egocentric RGB/depth, Manus glove tracking, IMU, timeseries, and videosOpen egocentric manipulation loader tests and sensor-fusion policy prototypingopen
HoyerChou Egocentric Bimanual Videos2026-01Hugging FaceYOTO/YOTO++ release with three raw first-person bimanual videos, 119 segmented demos, 241 rows, MANO hand meshes, 3D keypoints, calibration, and camera/world transformsOne-shot bimanual manipulation and hand-trajectory extraction from human videoopen
Egocentric Specialties2026-01Hugging FaceAuto-gated skilled-jewelry manufacturing data card describing synchronized egocentric and top-down footage for diamond setting and specialty workflowsSkilled industrial manipulation and fine-grained manufacturing demonstrationsrequest
X-Humanoid2025-12arXivApplies human-to-humanoid video translation to 60 hours of Ego-Exo4D, releasing 3.6M robotized humanoid framesHumanoid video/world-model data generationwatch
DreamTacVLA2025-12arXivGrounds VLA policies in contact physics using tactile images, wrist-camera local vision, and third-person macro visionContact-rich VLA manipulationwatch
TacThru-UMI2025-12arXivSee-through-skin tactile-visual sensing plus a UMI-style imitation-learning framework for multimodal manipulationTactile-visual robot policy learningwatch
Mitty2025-12arXivDiffusion Transformer for translating human demonstrations into robot-execution videos, with synthesis from large egocentric datasetsHuman-to-robot video generationwatch
World in Your Hands / WiYH2025-12arXivReported 1,000-hour egocentric manipulation ecosystem with multiview/depth/hand/wrist signalsVLA, manipulation representation learningwatch
Hoi!2025-12arXiv3,048 force-grounded articulated-manipulation sequences over human, wrist-camera, UMI, and Hoi-gripper embodimentsForce/contact-grounded cross-view manipulationwatch
METIS2025-11arXivDexterous VLA pretrained on multi-source egocentric human and robotic data unified under one action spaceDexterous egocentric-VLA pretrainingwatch
Let Me Show You / RfV2025-11IROS 2025Retrieves egocentric human demonstrations and extracts affordance masks plus hand trajectories for robot policiesRetrieval-from-video robot manipulationwatch
UMIGen2025-11arXivCloud-UMI plus visibility-aware egocentric point-cloud generation3D-aware cross-embodiment imitation learningwatch
EgoMI2025-10arXivEgocentric manipulation interface capturing synchronized end-effector and active-head trajectories for semi-humanoid robotsActive-vision whole-body manipulationwatch
FARM2025-10arXivForce-aware diffusion policy using a modified tactile UMI gripper and matched actuated deployment gripperForce- and tactile-grounded manipulationwatch
UMI-on-Air2025-10arXivEmbodiment-aware diffusion policy adapting handheld UMI demonstrations to aerial manipulatorsUMI transfer under embodiment dynamicswatch
EmbodiSwap2025-10arXivPhotorealistic robot overlays over ego-centric human video for synthetic robot imitation datasetsZero-shot human-video-to-robot imitationwatch
Hand-VLA Pretraining2025-10arXivConverts unscripted real-life egocentric hand videos into 1M VLA-aligned action-language episodesRobot pretraining from human activity videowatch
ActiveUMI2025-10arXivPortable VR/UMI-style system capturing active egocentric perception for bimanual manipulationActive-perception UMI-style data collectionwatch
FastUMI-100K2025-10arXiv100K+ UMI-style trajectories across 54 tasks in LeRobot format with multi-view wrist fisheye imagesLarge-scale UMI-style robot policy trainingopen
BiNoMaP2025-09arXivCategory-level bimanual non-prehensile primitives learned from egocentric hand trajectories and transferred across two dual-arm platformsContact-rich bimanual skill extraction from human demonstrationswatch
EMMA2025-09IEEE RA-L 2025Co-trains egocentric human mobile-manipulation data with static robot data to avoid mobile teleoperation bottlenecksScalable mobile manipulation from human datawatch
OpenEgo2025-09arXiv1,107 hours unified across six public egocentric datasets, 290 manipulation tasks, hand-pose/action primitivesStandardized dexterous manipulation pretraining and evaluationwatch
MV-UMI2025-09arXivMulti-view UMI adds third-person context to wrist egocentric observationsBroad-scene context for cross-embodiment manipulationwatch
InterVLA2025-08ICCV 202511.4 hours and 1.2M frames with 2 ego and 5 exo views, human/object motions, and verbal commandsVision-language-action and motion estimationwatch
EgoVLA2025-07arXivVLA training from egocentric human videos plus robot fine-tuningHuman-video-to-robot policy transferwatch
EgoDex2025-05ICLR 2026829 hours of Apple Vision Pro egocentric video, 194 tabletop tasks, 3D hand/finger trackingDexterous manipulation, imitation learning, human-to-robot hand trajectory predictionwatch
TASTE-Rob2025-03CVPR 2025100,856 task-oriented ego-centric hand-object interaction videos aligned with language instructionsVideo generation and robot imitation learningwatch
Kaiwu2025-03IEEE RA-L 202511,664 synchronized assembling-action instances with first-person video, gaze, hand motion, pressure, audio, EMG, mocap, and robot streamsMultimodal robot learning and HRIwatch
Humanoid-VLA2025-02arXivCombines language, egocentric scene perception, and motion control for universal humanoid controlHumanoid VLA and motion generationwatch
Immersive Social Interaction with VR and LLM-Assisted Humanoids2024-11Humanoids 2024 WorkshopApple Vision Pro teleoperation with egocentric robot feedback, voice locomotion, wrist/finger retargeting, and RGB, gaze, command, joint-state, and hand-motion recording on a Unitree H1Immersive humanoid teleoperation and multimodal demonstration capturewatch
EgoMimic2024-10ICRA 2025Project Aria human videos with 3D hand tracking and robot co-training for unified imitation learningHuman-to-robot imitation from ego videowatch
FastUMI2024-09arXivUMI redesign reporting 10K+ real-world trajectories across 22 everyday tasksFaster hardware-independent UMI-style collectionwatch
R+X2024-07ICRA 2025Retrieves executable skills from long unlabeled first-person videos of everyday human tasksRobot imitation from everyday first-person videowatch
UMI on Legs2024-07arXivCombines handheld UMI task demonstrations with simulation-trained whole-body controllers for quadruped manipulationMobile cross-embodiment UMI transferwatch
EgoPAT3D / EgoPAT3Dv22024-03ICRA 20241M+ RGB-D/IMU frames in the original task; v2 expands egocentric 3D action-target predictionHuman-robot interaction, 3D target anticipation, manipulation safetyopen
Universal Manipulation Interface / UMI2024-02RSS 2024Hand-held GoPro gripper interface and policy stack for in-the-wild robot teachingWrist-view robot-free demonstrations and cross-embodiment policy transferopen
Dobb-E / HoNY2023-11arXiv13 hours, 22 homes, 5,620 trajectories, and 1.5M RGB-D frames collected with an iPhone-based StickHousehold robot learning from handheld in-home demonstrationspartial
VIP2022-09ICLR 2023Value-Implicit Pre-Training learns visual rewards and representations from Ego4D human video for few-shot robot manipulationEgo4D-to-robot reward learningopen
R3M2022-03CoRL 2022Robot-manipulation visual representation pretrained on Ego4D with time-contrastive and video-language objectivesEgo4D-to-robot representation learningopen

Hands, Objects, Pose, Tracking, and 3D

Fine-grained hand, object, contact, and 3D-pose datasets, including emerging event-camera and EMG sensing.

ResourceReleasedVenueScale / signalBest forStatus
HOPE2026-08arXivPredicts temporally evolving per-vertex hand pressure and contact from monocular hand-centric video by unifying tactile-glove, planar-pressure, and contact labelsDynamic hand-object pressure and contact estimationwatch
EgoGVAE2026-07ECCV 2026Guided variational autoencoder reconstructing full-body meshes from head pose with one-step sampling and reported 50x+ faster inference than diffusion alternativesFast head-conditioned ego-body mesh recoverywatch
ReViV2026-07ECCV 2026Open unified feed-forward reconstruction of camera trajectory, gaze, body, hands, depth, and view dynamics from one monocular egocentric RGB streamHolistic and efficient viewer-plus-scene 4D reconstructionopen
Egocentric Hand Benchmark Annotations2026-07Hugging Face100,427 public annotation rows combining 95,002 EgoDaily hand boxes and 5,425 HOT3D virtual hand-crop camera records; source RGB is not redistributedHand detection, crop-camera geometry, and derived benchmark evaluationpartial
EgoExoMoCap2026-07ECCV 2026Distributed mocap from two or more smart-glasses wearers, fusing head/wrist tracking with context-aware image features on two in-the-wild datasetsLightweight ego-exo body-motion reconstructionwatch
Exo2EgoPose2026-07ACM MM 2026Vision-language-guided egocentric 3D hand-pose forecasting using reconstructed exocentric demonstrations, evaluated on AssemblyHands, Ego-Exo4D, EgoMe-pose, and CALVINHand forecasting and human-to-robot transferwatch
EgoObjects-Camera2026-07Hugging FacePublic camera-parameter annotations for 241,554 EgoObjects images across 49 shards, including roll, pitch, vertical field of view, and radial distortion; reuse terms are undeclaredCamera-aware egocentric modeling and dataset analysispartial
TSR-Ego2026-07arXivCausal temporal stereo refinement with fisheye deformable cross-attention for online egocentric 3D body pose on UnrealEgo2 and UnrealEgo-RWHead-mounted stereo pose recovery under occlusion and truncationwatch
CI-HOI / DEHOI2026-07arXivCue-isolated HOI evaluation and inpainted testbed that separate hand and object evidence to expose contextual shortcuts in egocentric VLMsDiagnosing hand- versus object-centric action understandingwatch
Whareformer2026-07ECCV 2026First learned OSNOM tracker, trained on 56 videos and tested on 260 long EPIC-KITCHENS-100, IT3DEgo, and HD-EPIC videos; MIT code, weights, features, and training data are liveLong-term 3D object tracking through occlusion and out-of-view periodsopen
HandsOnWorld2026-07arXivCamera-disentangled hand-controlled egocentric video generation with EgoVid-Pro: 103K clips and roughly 12M frames of protagonist-only hand trajectoriesUnconstrained egocentric HOI video generationwatch
SAGE Action-Gaze2026-07arXivUnified action-gaze recognition and anticipation framework linking HOI, gaze, egocentric close-up views, and exocentric contextJoint gaze/action anticipation for human behavior understandingwatch
Ego3DLM / Ego-Human Motion Prediction2026-07ECCV 20263D-aware LLM that jointly predicts past/future human pose and narration from egocentric video, three-point tracking, and scene features on NymeriaScene-grounded egocentric human-motion forecastingwatch
Humanola Egocentric Hand-Pose Sample2026-07Hugging FaceLeRobot v2.1 sample with 9 episodes, 48,272 frames, four synchronized cameras, 200 Hz IMU, 3D hands, poses, overlays, and Rerun logsOpen hand-pose and sensor-fusion loader tests for manipulation pipelinesopen
InterPet4D2026-06Hugging Face6.8M-frame human-dog capture; public 10.7 GB v1 provides 227 headset clips with audio and aligned SMPL, MANO, pet skeleton, SMAL, and MERT featuresMultimodal 4D human-pet interaction, motion generation, and pose analysisopen
ViDiHand2026-06arXiv4D two-hand motion reconstruction from full-frame egocentric video using video-diffusion representations; evaluated on ARCTIC, HOT3D, and HOI4DEgocentric hand-motion reconstruction without detector/test-time optimizationwatch
HT-Bench / HandTouch2026-06arXiv10M RGB frames and 7.8M tactile frames across 226 dexterous full-hand tasksVision-tactile representation learning for dexterous manipulationwatch
EPIC-Contact / HOPformer2026-06ECCV 20262.3K EPIC-KITCHENS clips and 62.3K frames with dense 3D hand-object contact, posed meshes, code, dataset, and checkpointsIn-the-wild egocentric 3D hand-object pose and contact estimationopen
GhostHandEgoNet2026-06IJPRAI 2026Lightweight egocentric hand-pose and hygiene-monitoring pipeline with 2.56M parameters for embedded deploymentEgocentric hand pose under occlusion and clinical hand-hygiene monitoringwatch
Hand-4DGS2026-06arXivFeed-forward 4D hand reconstruction from egocentric video with 3D Gaussian Splatting, evaluated on H2O and ARCTICFast egocentric 4D hand reconstructionwatch
A multimodal RGB/events FPV hand dataset2026-06arXivSynthetic event-based first-person hand detection from EgoHands plus v2eEvent/RGB hand detection benchmarkingwatch
SIGNIQ Egocentric Hand-Pose Sample2026-06Hugging FaceFive short head-mounted GoPro manipulation clips with Label Studio import metadata for hand-pose and action-segment annotationHand-pose annotation and manipulation-labeling workflow testsopen
Leveraging Synthetic Data for Egocentric HOI Detection2026-05IJCV 2026Synthetic-data augmentation study for hand-object interaction detection on VISOR, EgoHOS, and ENIGMA-51HOI detection when real labels are scarcewatch
Causal-Inspired Fourier Egocentric Action Recognition2026-05IEEE TCSVT 2026Causal-inspired Fourier representation learning for wearable IMUs and egocentric action recognitionWearable-sensor action recognitionwatch
EgoSMPLX2026-05ICIP 2026Prior-guided whole-body human mesh recovery from monocular head-mounted images, with public code and train/test pseudo-ground-truth packagesEgocentric whole-body mesh recoverypartial
DexGloveHOI2026-05arXiv100K+ synchronized egocentric vision and on-glove IMU samples with marker-based mocap 3D hand-pose ground truth across dexterous daily manipulationVision-IMU fusion for 3D hand trackingwatch
EggHand2026-05CVPR 2026 FindingsMultimodal foundation model for egocentric hand-pose forecasting over EgoExo4D-style video-language and action signalsHand-pose forecasting and VLA action decodingwatch
Map-Mono-Ego2026-05ICIP 2026Map-grounded global human-pose estimation from monocular egocentric video, with AIST-Living paired ego video and scanned environmentsScene-aware egocentric body-pose localizationwatch
EgoForce Motion2026-05arXivOnline full-body motion reconstruction from sparse egocentric observations using diffusion forcingReal-time egocentric body-motion reconstructionwatch
EgoEMG2026-05arXiv41 participants, bilateral EMG, IMU, RGB, external RGB-D, mocap hand labelsEMG plus egocentric vision hand pose estimationwatch
EgoTouch / TouchAnything2026-05arXiv302 tasks, 4,530 episodes with egocentric + dual wrist cameras, bimanual 3D hand pose, and dense tactile pressure mapsTactile estimation and contact modeling from egocentric videowatch
EgoEVHands2026-05arXivStereo event-camera egocentric hand dataset with 5,419 sequences and 3D/2D keypointsEvent-based bimanual hand pose and gesture recognitionwatch
TouchMoment2026-04CVPR 2026 Findings4,021 egocentric videos, 8,456 annotated hand-object contact momentsPrecise contact moment detectionwatch
TAIHRI2026-04arXivTask-aware VLM for localizing task-relevant 3D human keypoints in close-range egocentric HRIMetric-scale human keypoints for robot interactionwatch
MASS2026-04arXivDeformable surfel splatting for high-fidelity 3D hand reconstruction and rendering from egocentric monocular videoEgocentric hand reconstruction and renderingwatch
E-3DPSM2026-04arXivEvent-driven continuous pose state machine for real-time egocentric 3D human-pose estimation from head-mounted devicesEvent-based egocentric body-pose estimationwatch
EgoFun3D2026-04arXiv271 egocentric videos with 3D geometry, part segmentation, articulation and function-template annotationsInteractive 3D object modeling from ego videowatch
FunRec2026-04CVPR 2026Functional 3D digital-twin reconstruction from in-the-wild egocentric RGB-D interaction videos with articulated parts and kinematicsInteractive scene reconstruction for simulation and robot learningwatch
DP-DeGauss2026-04ICASSP 2026Dynamic probabilistic Gaussian decomposition for egocentric 4D scene reconstruction, separating background, hands, and objectsFirst-person 4D interaction reconstructionwatch
Towards Egocentric 3D Hand Pose Estimation in Unseen Domains2026-03WACV 2026Egocentric 3D hand-pose estimation study focused on unseen-domain generalizationRobust first-person hand tracking outside training domainswatch
Controllable Egocentric Video Generation2026-03ECCV 2026Occlusion-aware sparse 3D hand-joint control for realistic egocentric hand-object video generation, with a 1M-clip annotation pipelineControllable egocentric HOI video generation and world modelswatch
SHOW3D2026-03CVPR 2026Open 20-hour, 2,137-scene release from 38 participants with eight exo and two Quest 3 ego views, 3D hand/object annotations, shape, and textIn-the-wild 3D hand-object reconstructionopen
EgoXtreme2026-03CVPR 2026Smart-glasses egocentric 6D object-pose dataset across industrial, sports, and rescue scenes with extreme motion blur, dynamic lighting, and occlusionRobust 6D object pose under extreme egocentric conditionswatch
AG-EgoPose2026-03arXivAttention-guided egocentric 3D human-pose estimation from fisheye camera input with dual motion/spatial streamsFisheye egocentric 3D pose estimationwatch
FEEL2026-03arXivOpen 26-trial, 700-episode benchmark with 39,497 Aria frames and synchronized gaze, hand, IMU, and Paxini grasp-force totals; the paper reports a larger full collectionContact-rich physical action and hand-object understandingopen
Person Identification from Egocentric HOI2026-02CCF TPCI 2026Uses 3D hand pose from egocentric human-object interactions as a person-identification signalIdentity and behavior analysis from first-person hand-object motionwatch
WHOLE2026-02arXivJoint world-grounded hand-object motion reconstruction from egocentric video using a generative HOI priorHolistic hand-object pose recoverywatch
AirGlove2026-02ICASSP 2026Egocentric 3D hand-tracking evaluation for gloved hands and sensing-glove appearance shiftsGloved-hand tracking for teleoperationwatch
Eva-3M / EvaPose2026-02CVPR 20263.0M+ egocentric HPE frames, including 435K keypoint-visibility labels, plus visibility-aware pose estimationVisibility-aware egocentric human posewatch
WristPP2026-02CHI 2026 submissionWrist-worn wide-FOV RGB system with a 133K-frame pose-pressure dataset from 20 subjectsMobile hand pose and pressure interactionwatch
Egocentric Affordance Dataset2026-02Hugging Face48,400 egocentric frames across household, workshop, factory, gardening, and healthcare scenes with Qwen3-VL affordance labelsAffordance and HOI supervision for embodied VLMs and robot learningopen
Industrial Egocentric HOI Detection2026-01ICIAP 2025 WorkshopsReal-time egocentric hand-object interaction detection for industrial workflowsIndustrial first-person HOI monitoringwatch
EgoGrasp2026-01arXivWorld-space open-vocabulary hand-object interaction reconstruction from dynamic egoview videosHOI reconstruction in world coordinateswatch
GlovEgo-HOI2026VISAPP 2026Synthetic-to-real egocentric HOI detection for gloved industrial hand-object scenariosIndustrial HOI detection under gloves and domain shiftwatch
OpenTouch2025-12arXiv5.1 hours of synchronized egocentric video-touch-pose data, 2,900 curated clips, text annotations, and retrieval/classification benchmarksFull-hand tactile grounding for real-world interactionwatch
ROHIT2025-12arXivReconstructs objects along hand-interaction timelines from HOT3D and EPIC-KITCHENS stable-grasp clipsObject pose propagation through hand contactwatch
EvHand-FPV2025-09arXivEvent-based first-person 3D hand-tracking dataset and lightweight wrist-ROI tracking frameworkLow-power FPV hand trackingwatch
THOR2025-07arXivThermal-guided adaptive RGB sampling for wearable hand-object monitoring, using about 3% of RGB frames while preserving activity segmentsLow-power longitudinal hand-object activity recognitionwatch
Egocentric HOI Detection2025-06Expert Systems with ApplicationsNew benchmark and method for detecting hand-object interactions in egocentric videoEgocentric hand-object interaction detectionopen
EgoBrain2025-06ICLR 2026Complete gated 1.6 TB release from 40 participants with synchronized 4K egocentric video, EEG, IMU, interval markers, and surveysBrain-signal plus first-person action understandingrequest
EventEgoHands2025-05ICIP 2025Event-based egocentric 3D hand mesh reconstruction benchmark over N-HOT3DLow-light and motion-blur hand reconstructionwatch
EgoEvGesture2025-03arXivFirst large-scale egocentric event-camera gesture-recognition dataset with a head-motion-robust 7M-param model (62.7% unseen-subject accuracy)Event-camera egocentric gesture recognitionwatch
EgoCast2024-12WACV 20253D pose forecasting from egocentric video and proprioceptive data on Ego-Exo4D / Aria Digital TwinFuture body-pose forecastingwatch
Ego3DT2024-10ACM MM 2024Zero-shot 3D reconstruction and tracking of all objects in ego-centric videos3D object tracking in first personwatch
EgoPressure2024-09CVPR 20255.0 hours, 21 participants, a moving egocentric camera plus 7 stationary RGB-D cameras, hand pose meshes, and fine-grained per-contact touch pressure (CVPR 2025)Hand pressure and pose from egocentric visionopen
AFF-ttention / STAformer2024-06ECCV 2024Affordance- and attention-aware short-term object-interaction anticipation from egocentric videoShort-term active-object anticipationwatch
HOT3D2024-06CVPR 2025833+ minutes, 3.7M+ images, Aria and Quest 3, gaze, point clouds, 3D hand/object/camera poses3D hand-object tracking for AR/VR manipulationopen
EMAG2024-05ECCV 2024 HANDS WorkshopEgo-motion-aware 2D hand forecasting for egocentric videosHand trajectory anticipationwatch
UnrealEgo2 / UnrealEgo-RW2024-01CVPR 2024Expanded stereo egocentric pose datasets from synthetic and real-world captureStereo egocentric human pose estimationwatch
POV-Surgery2023-07MICCAI 2023Egocentric surgical video for hand and tool pose estimation during operating-room activitiesSurgical hand-tool pose and interaction understandingopen
EgoHumans2023-05ICCV 2023125K+ egocentric images for in-the-wild multi-human tracking, 2D/3D pose, and mesh recoveryEgocentric multi-human perception and trackingpartial
AssemblyHands2023-04CVPR 20233.0M 3D hand-pose annotations from Assembly101, including 490K egocentric imagesEgocentric 3D hand pose and pose-aware action recognitionbenchmark
EgoTracks2023-01NeurIPS 2023Long-term egocentric visual object tracking annotations from Ego4DTracking, re-detection, embodied object persistenceopen
EgoGTA / EgoPW-Scene2022-12CVPR 2023Synthetic EgoGTA plus EgoPW-based in-the-wild scene-aware pose dataScene-aware egocentric 3D human pose estimationpartial
EgoHOS2022-08ECCV 202211,243 egocentric images with detailed hand/object/contact segmentationHand-object segmentation and contact-aware foreground modelingopen
UnrealEgo2022-08ECCV 2022Synthetic stereo eyeglass egocentric dataset with 3D human poseRobust egocentric 3D human motion captureopen
Ego2HandsPose2022-06WACV 2024Extension of Ego2Hands with 3D global hand pose annotationsTwo-hand RGB 3D pose in global coordinatesopen
ARCTIC2022-04CVPR 20232.1M frames with bimanual hand/object meshes, articulated objects, and contactDexterous bimanual manipulation, contact, reconstructionrequest
GIMO2022-04ECCV 2022Ego-centric views, gaze, scene scans, and high-quality body pose sequencesGaze-informed human motion prediction in scene contextopen
HOI4D2022-03CVPR 20222.4M RGB-D egocentric frames, 4,000+ sequences, 800 objects, 16 categories, 610 rooms4D human-object interaction, segmentation, object pose tracking, action segmentationopen
EgoBody2021-12ECCV 2022HoloLens2 egocentric RGB/depth/eye/head/hand data with 3D body pose and shapeEgocentric human pose, shape, motion, social interactionopen
TREK-1502021-08project page150 egocentric tracking sequences derived from EPIC-KITCHENSSingle-object tracking in egocentric videoopen
H2O: Two Hands and Objects2021-04ICCV 2021Synchronized multi-view RGB-D, two-hand 3D pose, 6D object pose, object meshes, camera posesFirst-person two-hand interaction recognition and pose-aware HOIopen
Ego2Hands2020-11arXivLarge-scale composited RGB egocentric two-hand segmentation/detectionRobust unconstrained hand segmentationopen
xR-EgoPose2019-07GitHubSynthetic egocentric body pose data for XR headset viewpoints3D human pose from head-mounted camerasopen
You2Me2019-04CVPR 2020Infers second-person body pose in egocentric video from first- and second-person interaction cuesSocial interaction and body-pose inferencewatch
H+O2019-04CVPR 2019 OralUnified egocentric RGB model for 3D hand pose, object pose, interactions, objects, and actions3D hand-object pose and interaction recognitionwatch
EgoReID Dataset2018-12arXiv900 identities, about 10.2K tracks, 176K detections from mobile first-person cameras with sensor metadataPerson re-identification from first-person videowatch
EgoYouTubeHands / HandOverFace2018-03CVPR 2018In-the-wild egocentric hand-segmentation study built around EgoYouTubeHands, HandOverFace, EGTEA, and EgoHands-style supervisionHand detection and segmentation beyond lab capturewatch
FPHA / First-Person Hand Action2017-04CVPR 2018100K+ RGB-D frames, 45 action classes, 26 objects, 3D hand/object poses3D hand pose and first-person hand action recognitionopen
EgoGesture2017-01IEEE TMM 201824K+ gesture samples, 3M RGB-D frames, 50 subjects, 83 static and dynamic gestures across six scenesFirst-person gesture recognition for wearable interactionopen
Left/Right Hand Segmentation2016-07CVIU 2016Classic first-person method that separates left/right hands and handles hand-to-hand occlusion with temporal superpixelsHand segmentation and wearable interactionwatch
EgoHands2015-12ICCV 201548 Google Glass videos, 4,800 annotated hand imagesHand detection and segmentationopen
BEOID2014-09BMVC 2014Gaze-tracked egocentric video of 8 users interacting with objects across 6 everyday locations (kitchen, workspace, printer, corridor, gym)Discovering task-relevant objects and interaction modesopen

Daily Life, Memory, Assistance, and QA

Long-horizon daily-life capture and the question-answering benchmarks that probe memory, reasoning, and assistance.

ResourceReleasedVenueScale / signalBest forStatus
EgoMonth2026-08arXiv300+ hours from 20 participants over 20-120 days and 1,443 human-crafted questions across 14 tasks; annotations, scripts, and eight sample videos are public while full raw video uses controlled accessMonth-scale episodic memory, spatial reasoning, and personal QApartial
SERUM2026-07COLM 2026Extracts compact action and intent state models from 61 egocentric videos spanning coding, cooking, physical activity, and daily lifeInterpretable user workflows, intent modeling, and personalized assistantswatch
Sidewalk Moments2026-07arXiv61 first-person city walks segmented into more than 50K ten-second clips with video, averaged-image, audio, and text representationsHuman-aligned urban engagement and multimodal temporal-compression studieswatch
Vinci2 / EgoServe2026-07ECCV 20263,000+ proactive-service instances across 11 categories and four memory horizons, with public annotations and the training-free EgoMemo agentDeciding when an egocentric assistant should intervene and grounding its response in memoryopen
VEGAS2026-07arXivTraining-free gaze-aware caption metric plus egocentric activities and instructional slides paired with synchronized gaze and reference captionsHuman-attention-aligned caption evaluation and retrievalwatch
CoMind2026-07ECCV 2026Dual-view collaborative-cooking dataset with two synchronized head-mounted cameras, two exo views, audio, gaze, 3D scene/object scans, and social/interaction annotations; download is marked coming soonSocial reasoning, joint attention, handover prediction, and collaborative assistancewatch
Imprint2026-07arXivOnline interaction-centric memory compression for seven-day egocentric QA on EgoLifeQAScalable long-horizon memory retrievalwatch
EgoGapBench2026-07arXivDiagnostic benchmark for whether models select actions from the camera-wearer's perspective in multi-agent scenesEgocentric perspective and action-selection evaluationopen
EgoSafetyBench2026-06arXiv1,200 egocentric robot-view safety scenarios with half-second labels and misleading in-scene text controlsStreaming VLM safety-guard evaluationwatch
BinaryTracking / GangnamLoop2026-06arXivSpatial QA and navigation benchmark for long egocentric robot routes, with GangnamLoop route data and the BinTrack implementationOpen spatial memory and navigation QAopen
Egocentric Pedestrian Crossing Intention2026-06arXivVLM-based VQA formulation for predicting pedestrian crossing intention from short egocentric traffic-safety clipsWearable traffic-safety intent reasoningwatch
Face-vs-Body HRI Tracking2026-06arXivCustom-annotated egocentric HRI dataset from a social robot, comparing face and body tracking under occlusion and engagement changesSocial-robot engagement trackingwatch
VL-MemKnG / WalkieKnowledgeT+2026-06arXivHybrid spatio-temporal knowledge graph plus segment memory for QA over long egocentric navigation trajectoriesNavigation memory, long-horizon evidence retrieval, spatial QAwatch
V-RAGBench / CARVE2026-06arXivQuery, evidence-chunk, answer triplets for decoupled retrieval/generation evaluation in long egocentric video RAGRetrieval-augmented long-video QAwatch
OVO-S-Bench2026-06arXiv1,680 spatial-reasoning questions over 348 continuous egocentric videos with query timestamps and evidence intervalsStreaming spatial intelligence, tracking, simulation, and allocentric mappingwatch
Astra2026-06arXivAgentic VLM spatial reasoning with a view-consistent world simulator that supplies imagined novel-view evidence from limited egocentric observationsWorld-model-augmented spatial reasoningwatch
VLESA2026-06arXivGoal-conditioned safety annotations over egocentric human-activity video for safety-aware assistanceSafety monitoring and situated activity assistanceopen
Causal-Plan-1M2026-06arXivReported million-scale corpus of explicit causal reasoning traces over egocentric videosCausal and planning-oriented ego reasoningwatch
NoRA2026-06arXiv1,420 first-person video clips for normative action reasoning with fact-reason-action support graphsGrounded reasonableness and safety-oriented action generationwatch
Pause and Think2026-06arXivReasoning-centric training data and benchmark for video-grounded assistive action suggestionsScene-grounded assistance, planning, and temporal consistencywatch
SuperMemory-VQA2026-06arXiv52.9 hours of AI-glasses activity with RGB, audio, gaze, IMU, SLAM, and 4,853 human-verified QA pairs across object/location/intent/scene memoryLong-horizon memory for AR assistantswatch
VISTA Daily Assistance2026-05arXivGenerative egocentric-video framework for reactive and proactive daily-assistance scenariosSynthetic training/evaluation for assistantswatch
EgoMemReason2026-05COLM 2026500 public questions over week-long EgoLife video with entity, event, and behavior memory, open evaluation code, and a leaderboardMemory-driven reasoning across sparse evidence over hours or daysopen
EgoBench2026-05arXiv1,045 egocentric-video-grounded interactive tasks with tools and simulated usersTool-using multimodal agents with dynamic interactionwatch
EyeCue2026-05arXivGaze-empowered egocentric video model plus CogDrive annotations for driver cognitive-distraction detectionGaze-context reasoning for safety and internal-state inferencewatch
Minerva-Ego2026-05arXivComplex egocentric visual-reasoning benchmark with multi-step multimodal questions, dense reasoning traces, and spatiotemporal object masksGrounded multi-step egocentric reasoningopen
EgoPro-Bench2026-05arXivPersonalized proactive-interaction benchmark with 2,400 evaluation videos, 12K+ training videos, and 12 domainsStreaming intent prediction and proactive assistance timingwatch
Pro2Assist2026-05arXivContinuous step-aware proactive assistance with multimodal egocentric perception and an AR-glasses testbedReal-time step-aware assistancewatch
Personal Visual Context Learning / Personal-VCL-Bench2026-05arXivSmart-glasses benchmark for using wearer-specific visual context at prompt time over continuous first-person streamsPersonalized visual memory for wearable assistantswatch
AwareLLM2026-05arXivProactive assistant using egocentric vision, pupillometry, gaze, posture, heart activity, and LLM reasoningPersonalized productivity and cognitive-state-aware assistancewatch
GazeMind / CogLoad-Bench2026-05arXivGaze-guided LLM agent and cognitive-load benchmark for smart glasses, reporting 152 participants, 40+ hours, and 10K+ annotationsCognitive-load-aware assistancewatch
Dietary Behavior Change Receptivity2026-05arXivPilot AIM-2 wearable-camera study using egocentric eating images to infer receptivity for just-in-time dietary interventionsNutrition and health-oriented lifeloggingwatch
EgoBabyVLM2026-05arXivBenchmarking cross-modal learning from naturalistic infant and adult egocentric video, including Machine-DevBench and challenge tracksDevelopmental egocentric VLM learningwatch
EgoIntrospect2026-05arXiv180 hours from 60 subjects with synchronized video, audio, gaze, motion, physiological signalsInternal-state reasoning, affect, intent, cognitive memory for wearable assistantswatch
EgoExoMem2026-05arXiv2.6K MCQs across synchronized ego-exo videosCross-view memory reasoningwatch
EgoCoT-Bench2026-05arXiv3,172 verifiable QA pairs over 351 videos with operation-centric rationale annotationsGrounded chain-of-thought and evidence consistencyopen
EgoStream2026-05arXiv2,250 questions across seven memory dimensions with Answer Validity Windows, expanded to 8,528 recall-conditioned evaluations over streams up to 45.3 hoursDiagnosing streaming episodic memorywatch
ChildLens2026-04Behavior Research Methods 2026109 hours of vest-mounted first-person video/audio from 62 children aged 3-5 at home, with five location and 14 activity classesNaturalistic child activity, temporal localization, and voice analysisrequest
MyEgo2026-04CVPR 2026541 long videos and 5K personalized questions about the camera wearer, belongings, activities, and pastPersonalized ego-grounding and long-range memory QAopen
EgoEsportsQA2026-04arXiv1,745 QA pairs from professional first-person shooter matches across three gamesHigh-speed first-person perception and tactical reasoning in virtual environmentswatch
EgoEverything2026-04arXiv5,000+ MCQ pairs over 100+ hours, with gaze-attention-inspired question generation for ARHuman-behavior-inspired long-context egocentric understandingwatch
LifeDialBench / EgoMem2026-04arXivReal egocentric video plus synthetic lifelog conversationsOnline temporal-causality evaluation for continuous lifelog memory systemswatch
EgoScreen-Emotion / ESE2026-04arXiv224 egocentric screen-view movie trailers with 28,667 confidence-aware emotion keyframe annotationsEmotion understanding from realistic first-person viewingwatch
Gaze-to-Guidance Assistants2026-04arXivGaze-grounded multimodal LLM assistance study using egocentric video and gaze overlays to infer reading difficultyCognitive-need-aware wearable assistancewatch
EgoTL2026-04arXivThink-aloud (say-before-act) egocentric capture with word-level spoken reasoning and metric-scale spatial annotations over 100+ household tasksLong-horizon reasoning for VLMs and world modelswatch
EgoSelf2026-04arXivPersonalized egocentric assistant framework with graph memoryPersonalization from long-term egocentric interaction memorywatch
MA-EgoQA2026-03arXiv1.7K questions over multiple long-horizon egocentric streamsMulti-agent egocentric memory reasoningopen
Ego2Web2026-03CVPR 2026Egocentric videos paired with web tasks requiring physical-scene understanding and online executionWeb agents grounded in first-person physical contextwatch
Egocentric Co-Pilot2026-03WWW 2026Web-native smart-glasses agent framework with temporal reasoning, context compression, speech/gaze intent, and streaming WebRTC/WebSocket pipelinesAssistive always-on egocentric web agentswatch
LifeEval2026-03arXiv4,075 QA pairs over continuous first-person streams for real-time task-oriented human-AI collaboration in daily lifeDaily-life multimodal assistance evaluationwatch
Gesture-Based Egocentric Video QA2026-03CVPR 2026Egocentric video QA grounded in the camera wearer's pointing and deictic gesturesGesture-grounded referential QAwatch
EgoIntent2026-03arXiv3,014 steps over 15 daily-life scenarios for local intent (What), global intent (Why), and next-step plan (Next)Step-level intent and anticipatory assistancewatch
SAW-Bench2026-02ICML 2026Open 3.30 GB benchmark with 786 Ray-Ban Meta smart-glasses videos and 2,071 QA pairs across six situated-awareness tasksFirst-person situated spatial reasoningopen
EgoGraph2026-02arXivDynamic temporal knowledge graph framework evaluated on EgoLifeQA and EgoR1-benchTraining-free ultra-long egocentric video memorywatch
EgoSound2026-02CVPR 20267,315 validated QA pairs across 900 Ego4D/EgoBlind videosSound perception, spatial localization, causal audio-visual reasoningwatch
SuperGlasses2026-02CVPR 2026 FindingsEvaluates VLMs as intelligent agents for AI smart glasses across realistic wearable assistant tasksSmart-glasses agent evaluationwatch
Gazeify Then Voiceify2026-01IUI 2026Displayless smart-glasses object referencing via gaze selection, mask generation, voice descriptions, and conversational correctionGaze-plus-voice physical object groundingwatch
Wearable Product Localization2026-01arXivWearable assistive shopping system combining detection, VLM guidance, spatialized sonification, and corrective feedbackProduct search and navigation for blind or low-vision userswatch
Egocentric Clinical Intent2026-01arXivBenchmark for egocentric clinical intent understanding by medical multimodal LLMs over first-person clinical procedure videoClinical intent reasoning for medical assistantswatch
WearVox2026-01arXivEgocentric multichannel voice-assistant benchmark for spoken interaction grounded in first-person audio-visual contextWearable voice-assistant evaluationwatch
Screen Detection from Egocentric Streams2026IEEE TMM 2026Multi-view vision-language approach for detecting screens in egocentric image streamsWearable screen understanding and assistive scene-text interfaceswatch
Ego-EXTRA2025-12WACV 202650 hours of expert-trainee procedural assistance dialogue, 15K+ VQA pairsEgocentric video-language assistants and expert feedbackopen
WearVQA2025-11NeurIPS 20252,520 image-question-answer triplets across 7 domains and 10 task types under occluded/low-light/blurry wearable capture (NeurIPS 2025)Smart-glasses VQA under real wearable conditionswatch
WAGIBench2025-10NeurIPS 2025 Spotlight29 hours, 348 participants, and 3,477 multimodal recordings for assistive wearable-agent goal inferenceInferring user goals from wearable contextwatch
TeleEgo2025-10arXivStreaming omni-modal benchmark with 3,291 QA items across memory, understanding, and cross-memory reasoningReal-time egocentric AI assistant evaluationwatch
CIVIL Lifelog Retrieval2025-10arXivCaptioning-integrated visual lifelog retrieval that describes wearable-camera images as first-person experiences before text-query matchingPersonal memory search and lifelog retrievalwatch
EgoNight2025-10ICLR 2026EgoNight-VQA (3,658 QA / 90 videos / 12 types) plus day-night retrieval and depth, with day-night aligned videoNighttime / low-light egocentric robustnessopen
EgoMemory2025-09OpenReview165,795 user-specific object annotations over 245 videos from 45 participantsMemory-augmented personalized retrievalwatch
HowToDIV2025-08arXiv507 conversations, 6,636 QA pairs, 24 hours of egocentric instructional task-assistance video clipsInstructional dialogue and procedural video reasoningwatch
EgoTrigger / HME-QA2025-08ISMAR 2025 / TVCGAudio-driven smart-glasses capture strategy plus 340 first-person HME-QA pairs from full-length Ego4D videosEnergy-efficient memory assistance and audio-triggered capturewatch
EgoCross2025-08AAAI 2026Cross-domain benchmark over surgery, industry, extreme sports, and animal perspectives; the EgoVis 2026 challenge drew 1,500+ submissions and 130+ participantsCross-domain egocentric QA generalizationwatch
ProMemAssist2025-07UIST 2025Smart-glasses proactive assistance via real-time working-memory modeling from multimodal wearable signalsTimely proactive assistancewatch
EgoPrivacy2025-06ICML 2025Benchmark for privacy leakage from first-person cameras across demographic, individual, and situational inferencePrivacy risk evaluation for ego videowatch
EASG-Bench2025-06ICCV 2025 WorkshopEgocentric action-scene-graph-based video QA benchmarkScene graph, temporal order, and relation-aware QAopen
Ego-R12025-06arXivEgo-CoTT-25K, Ego-QA-4.4K, and Ego-R1 Bench for tool-augmented ultra-long egocentric QAChain-of-tool-thought reasoning over days/week-long videoswatch
ExAct2025-06arXivVideo-language benchmark for expert action analysis and feedback over skilled activityExpert feedback and skill assessmentopen
EgoTextVQA2025-06CVPR 2025Egocentric scene-text-aware video QA across housekeeping and driving scenes (CVPR 2025)Reading and reasoning over first-person scene textopen
Week-Long First-Person Video Privacy Probe2025-04arXiv54-hour wearable-camera study of what foundation models infer from week-long first-person videoPersonal privacy and lifelogging riskwatch
EgoLog2025-04arXivAudio-IMU daily-log system for contextual and spatial activity logging with ubiquitous wearablesFine-grained lifeloggingwatch
EgoLife2025-03CVPR 2025300 hours from six participants over one week with Meta Aria, third-person references, and EgoLifeQALong-term daily-life memory, personalized assistants, ultra-long QApartial
EgoToM2025-03arXivTheory-of-mind QA over Ego4D-style egocentric videosGoals, beliefs, and next-action reasoning for camera wearerswatch
X-LeBench2025-01EMNLP 2025432 simulated life logs over Ego4D-like footage, spanning 23 minutes to 16.4 hoursExtremely long egocentric video understandingwatch
VidEgoThink2024-10ICLR 2025Ego4D-based benchmark for video QA, hierarchy planning, visual grounding, and reward modelingEmbodied egocentric video understandingbenchmark
Egocentric 360 VideoQA for Visual Impairment2024-05arXiv360-degree egocentric wearable-camera VideoQA dataset for multiple real-world obstacles faced by people with visual impairmentsAssistive VideoQA and full-surround wearable perceptionwatch
EgoThink2023-11CVPR 2024First-person VQA benchmark covering six capability groups and twelve dimensionsFirst-person perspective reasoning for VLMsbenchmark
EgoSchema2023-08NeurIPS 20235K+ curated multiple-choice QA pairs over 250+ hours from Ego4D clipsVery long-form video-language understandingopen
EgoTaskQA2022-10NeurIPS 2022Diagnostic QA benchmark for task dependencies, effects, intents, beliefs, and counterfactuals in ego videoTask-step reasoning and procedural QAbenchmark
AssistQ2022-03ECCV 2022531 question-answer samples from 100 newly filmed instructional videosAssistance-oriented video QA and affordance-centric task completionopen
MMAC Captions2021-09ACM MM 2021Sensor-augmented egocentric-video captioning data around CMU-MMAC-style multimodal activity streamsVideo captioning with RGB, audio, IMU, and textwatch
EgoVQA2019-10ICCV 2019600+ QA pairs over egocentric videos; an early first-person VideoQA benchmarkClassic egocentric VideoQAopen
HARMONIC2018-07IJRR 2022Assistive eating HRI data with head-mounted ego video, gaze, EMG, joystick, robot state, and third-person stereo viewsIntention prediction and shared-autonomy assistancewatch
First-Person Stories2017-07ICIAP 2017 Workshop45K+ egocentric photo-stream images labeled for lifestyle patterns such as eating, socializing, and sedentary behaviorLifelogging and daily-life behavior analysiswatch

Action, Procedure, Lifelogging, and Classic FPV

Procedural and activity datasets, from modern industrial assembly to the classic first-person-vision benchmarks that started the field.

ResourceReleasedVenueScale / signalBest forStatus
EventKitchen2026-08ECCV 20265.5 hours from 10 participants in 13 kitchens with stereo events, synchronized RGB/depth/IMU, 10,762 action segments, and 13,482 boxesNeuromorphic action recognition, object detection, and stereo depthwatch
Wearable Gait MoCap with Shank-Mounted Ego Cameras2026-07Scientific Data 2026Wearable motion-capture dataset for gait analysis using IMUs and shank-mounted egocentric camerasWearable gait analysis and motion-sensor fusionwatch
Poseidon Egocentric Activity Video Samples2026-06Hugging FaceTwo auto-gated 1080p, 30 fps everyday activity clips with task, scene, technical, feature, and PII metadataLightweight activity-footage and capture-quality reviewrequest
OR-Action2026-06arXivFine-grained multi-role operating-room action benchmark over public ego-exocentric OR video and scene-graph state changesSurgical workflow and temporal action understandingwatch
SkillSpotter2026-06ECCV 2026Pose-aware multi-view skilled-action detection and grading over Ego-Exo4D proficiency demonstrations, transferring to HoloAssistEgo-exo skill detection, grading, and coachingopen
TimeScribe Egocentric Annotations2026-06Hugging FaceManual dense-caption annotations for eight egocentric videos with frame-aligned micro-events, Chinese captions, English mirrors where available, source clips, and JSON manifestsHuman captioning gold data for hand-action and procedure understandingrequest
Egocentric Nursing Competency Assessment2026-05CVPR 2026 Workshop3.8 hours of egocentric nursing-simulation video with 493 annotated actions tied to instructor-rated competencyMedical skill assessmentwatch
IMPACT-HOI2026-05arXivMixed-initiative tool for constructing onset-anchored HOI event graphs in egocentric procedural videoStructured HOI annotation for robot learningwatch
EgoMAGIC2026-04arXiv3,355 egocentric field-medicine videos over 50 tasks from a head-mounted stereo camera with audio; 1.95M labels, 124 objects, action-detection challenge (Zenodo)Field-medicine perception, action and object detectionopen
PIE-V2026-04arXivMistake-aware procedural egocentric-video benchmark injecting plausible mistakes and recovery corrections across Ego-Exo4D scenariosProcedural mistake detection and recovery reasoningwatch
IMPACT2026-04arXivEgo-exo RGB-D industrial assembly dataset with bimanual, state, and anomaly annotationsIndustrial assembly, procedural state tracking, anomaly recoverywatch
ENIGMA-3602026-03arXivEgo-exo dataset for human behavior understanding in industrial scenarios with 360-degree and egocentric captureIndustrial ego-exo behavior understandingwatch
SAVA-X2026-03CVPR 2026Ego-to-exo imitation-error detection over asynchronous, length-mismatched egocentric and exocentric videos, evaluated with EgoMeCross-view imitation-error detectionwatch
Motion Focus Recognition2026-01arXivReal-time motion-focus recognition for fast-moving egocentric video using camera-pose foundation-model features and sliding-batch inferenceSports/fast-motion FPV intent and edge analysiswatch
SmartSeg2026Pervasive and Mobile Computing 2026Non-parametric temporal segmentation method for wearable-camera video streamsLifelog event segmentation and procedure understandingwatch
PEDESTRIAN2025-12arXiv340 first-person pavement videos covering 29 urban sidewalk obstacle types, with deep-learning detection baselinesPedestrian-safety obstacle detection from first-person videowatch
Mistake Attribution / MATT2025-11CVPR 2026Fine-grained mistake attribution over EPIC-KITCHENS-M and Ego4D-M with semantic, temporal, and spatial labelsProcedural mistake understandingwatch
SkillSight2025-11arXivFirst-person skill assessment using video+gaze teacher models distilled to a low-power gaze-only studentGaze-based skill learningwatch
IndEgo2025-11NeurIPS 2025~197h egocentric (plus ~97h exocentric) industrial collaborative work over assembly, logistics, inspection, and repair; gaze, narration, sound, motion, hand pose, point cloudsIndustrial egocentric assistants and procedure understandingopen
EgoEMS2025-11AAAI 2026High-fidelity multimodal egocentric data for cognitive assistance in emergency medical services, capturing time-critical team actionsReal-time medical procedural assistancewatch
EgoExOR2025-05arXivEgo-exo operating-room dataset with wearable RGB/gaze/hand/audio, RGB-D cameras, ultrasound, and scene graphsSurgical activity understandingwatch
Physical Activity VLM Annotation2025-05arXivVLM and discriminative-model study for reducing annotation burden in free-living wearable-camera physical-activity datasetsHealth-oriented activity labeling from lifelog imageswatch
LSC-ADL2025-04ACM MM 2025ADL annotations over lifelogging data generated with clustering plus human reviewActivity-aware lifelog retrievalopen
EgoSurgery2025-03Healthcare Technology LettersEgoSurgery-Phase (surgical phase recognition) and EgoSurgery-HTS (pixel-wise hand-tool segmentation of 14 tools) from egocentric open-surgery video (MICCAI 2024)Surgical phase, hand, and tool understandingopen
EgoMe2025-01arXivReal-world following-me dataset pairing exocentric demonstrations with egocentric imitation across everyday tasksEgocentric imitation and cross-view followingwatch
Audio-Narrated FPV Domain Generalization2024-09arXivMultimodal first-person action-recognition framework using audio narrations and audio-visual consistency to improve ARGO1M domain generalizationRobust first-person action recognition across environmentswatch
PARSE-Ego4D2024-06arXivEgo4D personal action recommendation annotations with 18K+ candidate suggestions and human preference evaluationAction recommendations for assistantswatch
Object-Aware Egocentric Online Action Detection2024-06CVPR 2024 EgoVis WorkshopObject-aware streaming action detection module for first-person video, validated on EPIC-KITCHENS-100Online action detectionwatch
EgoExo-Fitness2024-06ECCV 2024Synchronized ego and exo fitness videos with keypoint verification, execution comments, and quality scores (ECCV 2024)Ego-exo full-body action quality and skill assessmentopen
AIM-2 Food Ingestion Environment2024-05arXivTwo-stage transfer-learning method for recognizing food-ingestion environment from AIM-2 egocentric wearable-camera imagesNutrition and free-living dietary context recognitionwatch
EgoPack / Backpack Full of Skills2024-03CVPR 2024Multi-task egocentric video understanding with portable task perspectives across downstream skillsMulti-skill egocentric understandingwatch
EgoExoLearn2024-03CVPR 2024120 hours of egocentric plus demonstration videos with gaze and annotationsBridging asynchronous ego and exo procedural activityopen
Visual Experience Dataset / VEDB2024-02Journal of Vision240+ hours egocentric video with gaze/head tracking in classic literatureLifelogging, attention modeling, visual experience statisticspartial
CaptainCook4D2023-12NeurIPS 2024 Datasets and BenchmarksProcedural-activity dataset for understanding execution errors, later reused in multimodal temporal-action-segmentation workProcedural errors and mistake-aware activity understandingwatch
IndustReal2023-10WACV 2024Industrial-like egocentric procedure-step recognition dataset with execution errorsIndustrial procedure recognition and error handlingwatch
EGOFALLS2023-09ICPR 2024Visual-audio egocentric dataset and benchmark for fall detectionWearable fall detection and safety monitoringwatch
WEAR2023-04IMWUT 202422 participants, 18 outdoor workouts, synchronized egocentric video and 3D acceleration across 11 locations (IMWUT 2024)Vision-plus-inertial outdoor activity recognitionopen
FT-HID2022-09GitHub90K+ RGB-D first- and third-person human interaction samples from 109 subjectsFPV/TPV aligned human interaction analysisopen
EgoProceL2022-07ECCV 202262 hours, 130 subjects, 16 procedural tasksKey-step localization and procedure learning from ego videosopen
KrishnaCam / OAK2021-08ICCV 2021Long-running Google Glass daily-life stream; OAK adds 17.5 hours of object annotationsContinual object detection, summarization, personal visual memorypartial
Home Action Genome / HOMAGE2021-05CVPR 202127 participants, synchronized ego and third-person views, 12 sensor types, hierarchical activity/action labelsCompositional multi-view home activity understandingopen
MECCANO2020-10WACV 2021Industrial-like motorbike model assembly with RGB/depth/gaze variantsEHOI, active object detection, anticipation, industrial proceduresopen
EgoK3602020-10ICIP 2020First-person 360-degree activity videos360-degree egocentric activity recognitionpartial
LEMMA2020-07ECCV 2020Multi-view multi-agent multi-task daily activities across 14 kitchens/living rooms with dense atomic-action and HOI labels (ECCV 2020)Multi-agent compositional activity understandingopen
EGO-CH2020-02Pattern Recognition Letters 202027+ hours, 70 subjects, two cultural sites, 26 environments, 200+ points of interestCultural heritage visitor behavior, localization, retrieval, preference predictionpartial
Charades-Ego2018-04CVPR 201868.8 hours paired first-person and third-person videos, 68K+ activity instancesEgo-exo domain transfer, action recognition, localization, captioningopen
GTEA2018-01project pageEarly first-person cooking/activity datasetClassic ego action recognition and object interactionopen
GTEA Gaze / EGTEA Gaze+2018-01project pageMeal-prep egocentric videos with gaze and action annotationsGaze-aware action recognition and hand-object attentionopen
Ego2Top2016-07ECCV 201650 top-view videos, 188 egocentric videos, 166K frames, and 100K body detectionsMatching camera wearers across egocentric and overhead viewswatch
Ego-Engagement2016-04ECCV 2016Egocentric video dataset and model for whether the wearer is engaged with people or objectsEngagement, attention, and object/person interaction cueswatch
Wrist-mounted ADL2015-11CVPR 2016Synchronized head and wrist wearable-camera daily activitiesComparing head- versus wrist-mounted first-person viewsopen
Ego-Object Discovery / EDUB2015-04arXiv4,912 egocentric daily-life photo-stream images from four users for unsupervised object discoveryLifelog object discovery and detectionopen
HUJI EgoSeg2014-06project pageLong egocentric videos for temporal segmentationEgocentric event segmentationpartial
DogCentric2014-01project pageDog-mounted first-person activity videosAnimal egocentric activity recognitionopen
JPL First-Person Interaction2013-01IEEEFirst-person videos of people interacting with a humanoid observerHuman interaction recognition from first personpartial
First-Person Social Interactions2012-06CVPR 2012Day-long head-mounted video of 8 subjects at a theme park, annotated for social interactions, roles, attention, and turn-takingFirst egocentric social-interaction datasetopen
ADL Dataset2012-06project pageUnscripted daily activity recordings with activity/object/hand annotations in classic literatureDaily living action recognition and object interactionpartial
UT Ego2012-06project pageLong daily egocentric videos in classic summarization workTemporal segmentation and summarizationpartial
CMU-MMAC2009-06CMU tech report 2009Multimodal kitchen-activity database: head-mounted egocentric video plus body IMUs, motion capture, and audio for 43 subjects cooking 5 recipesOne of the first egocentric activity datasetsopen

Project Aria, AR/VR, and 3D Scene Resources

AR-glasses and headset data with gaze, SLAM, and digital-twin annotations for scene-level perception.

ResourceReleasedVenueScale / signalBest forStatus
SoniSpeech2026-08arXivOpen 34-hour eyewear dataset with 18,000 voiced/silent utterances, synchronized ultrasound, audio, and frontal video, covering 5,356 wordsOpen-vocabulary wearable silent-speech interfacesopen
Assistant Placement Aria2026-08ICRA 2026Synthetic and real scenes with 2D images, 3D point clouds, text, and three virtual-placement tasks under global, local, and human constraintsEgocentric spatial placement assistance and AR scene reasoningwatch
CRAFT Wearable Creative AI2026-07arXivSmart-glasses creative-AI probe informed by nine interviews, co-design with 16 participants, and 24 real-world sessions with eight writersContext-aware wearable creativity and in-situ human-AI collaborationwatch
SHARE User-Centric AR SLAM2026-07arXivUser-prioritized edge SLAM for commercial AR headsets and a ground robot, reporting 13.22 ms AR latency and sub-2-centimeter trackingResponsive AR interfaces in shared human-robot workspaceswatch
Ego6D Nymeria Features2026-07Hugging FacePublic Nymeria-derived 3D scene voxels, 20 Hz head 6-DoF windows over 149 scenes, and world-frame SMPL body pose over 148 scenes; terms are research-only and incompleteJoint 3D scene, localization, and body-pose experimentspartial
HeadRoom2026-07arXivEdge smart-glasses pipeline estimates visual and auditory channel availability from egocentric video/audio; a 25-participant study tests adaptive notification routingPerceptual-load-aware wearable assistancewatch
Future-Privileged Causal Ego Gaze2026-07arXivCausal egocentric gaze-estimation study showing future-aware training improves online models on EGTEA Gaze+ and Ego4DReal-time gaze modeling for AR/wearableswatch
EgoMed-IEMIS2026-06arXiv523 smart-glasses videos and 173,657 frames with referred medical-target masks across five imaging modalities and five viewing scenesInteractive egocentric medical segmentation and assistancepartial
AmbientEye2026-06arXivPassive-IR smart-glasses pupil-segmentation dataset for natural ambient illumination without active IR lightingEnergy-efficient eye tracking for all-day ARwatch
Single-View Mesh Rotation Stress Test2026-06arXivControlled camera-rotation stress test for single-view mesh reconstruction on Aria Digital Twin and Franka wrist-camera sequencesRobust 3D reconstruction under ego/wrist camera motionwatch
EPIC Efficient Egocentric Perception2026-06arXivSmart-AR-glasses perception framework using gaze, pose, and inertial signals to reduce high-resolution egocentric memory and energy costEfficient always-on AR perceptionwatch
EgoTraj2026-05arXiv75 Meta Quest Pro navigation sequences with RGB, head pose, gaze, scene labelsEgocentric human trajectory prediction and assistive navigationopen
RoSHI2026-04arXivWearable suit fusing sparse IMUs with Project Aria to estimate global 3D pose and body shape for in-the-wild human dataRobot-oriented human data capturewatch
GIST2026-04arXivSemantic-topology pipeline converting mobile point clouds into lightweight maps for search, localization, zone classification, and egocentric routing instructionsAssistive navigation in cluttered real spaceswatch
VueBuds2026-03arXivCamera-integrated wireless earbuds streaming low-power binocular views to on-device VLMsEarbud-scale wearable visual intelligencewatch
Pandora Articulated 3D Scene Graphs2026-03BMVC 2025Recovers articulated 3D scene graphs from Project Aria egocentric exploration for robot mappingArticulation-aware scene graphswatch
EgoPoseVR2026-02IEEE VR 2026 / IEEE TVCGOpen 388 GB synthetic RGB-D/HMD corpus with 18,235 motion clips, seven VR scenes, SMPL body parameters, and pose labelsSensor-free full-body VR pose estimationopen
ARGaze2026-02arXivAutoregressive transformer for causal online gaze estimation from first-person video and bounded recent gaze contextStreaming gaze estimation for AR/wearableswatch
EgoCampus2025-12arXivEgocentric pedestrian eye-gaze dataset and model over campus walking routes with synchronized first-person video and gazePedestrian gaze and trajectory predictionwatch
Eyes on Target2025-11RAAI 2025Depth-aware gaze-guided object detection for egocentric videoGaze-aware object perceptionwatch
egoEMOTION2025-10NeurIPS 202550+ hours from 43 participants with Project Aria video, gaze, PPG, IMU, emotion, and personality labelsAffect and personality from wearable capturewatch
Aria Gen 2 Pilot Dataset2025-10arXivIncremental multimodal daily-activity release captured with Aria Gen 2 glasses across cleaning, cooking, eating, playing, and outdoor walkingAria Gen 2 wearable sensingopen
LookOut / Aria Navigation Dataset2025-08ICCV 20254 hours of Project Aria real-world navigation recordings for future 6D head-pose trajectory predictionHumanoid and assistive egocentric navigationwatch
Aria STDR2025-07arXivProject Aria egocentric scene-text dataset studying lighting, distance, resolution, and gaze-guided OCRScene text recognition under wearable capturewatch
Photoreal Scene Reconstruction from an Egocentric Device2025-06SIGGRAPH 2025Project Aria egocentric-device reconstruction study improving pose and exposure treatment for pixel-accurate photoreal HDR scene reconstructionPhotoreal AR scene reconstructionwatch
Egocentric Event-Based Ping Pong2025-06CVPRW 2025Event-camera first-person table-tennis system and dataset for high-speed ball trajectory prediction under low latency and motion blurEvent-based wearable sports perceptionwatch
Reading in the Wild2025-05NeurIPS 2025100 hours of Project Aria reading/non-reading video with RGB, gaze, head pose, and reading-type labelsReading recognition for always-on smart glasseswatch
Digital Twin Catalog / DTC2025-04CVPR 20252,000 scanned objects with DSLR and egocentric AR-glasses image sequencesDigital-twin object reconstruction and evaluationopen
Head+Cane Camera Navigation2025-04arXivSynchronized head- and cane-mounted camera comparison for blind last-mile navigation, SLAM, and 3D reconstruction across five real environmentsHybrid wearable/cane sensing for assistive navigationwatch
egoPPG2025-02ICCV 2025Heart-rate estimation from eye-tracking cameras in egocentric systems to add physiological-state awareness to downstream wearer-behavior tasksPhysiology-aware egocentric perceptionwatch
HEADS-UP2024-09arXivHead-mounted egocentric dataset for pedestrian trajectory prediction and collision-risk assessment in blind-assistance systemsTrajectory prediction for blind assistancewatch
Smart-Glasses Engagement Prediction2024-09arXiv34-participant smart-glasses dyadic-conversation dataset with self-reported engagement ratings and LLM multimodal-transcript fusionSocial interaction and engagement reasoning from wearable capturewatch
Nymeria2024-06ECCV 2024300 hours, 264 participants, 50 locations, Aria, motion, language, observer viewEgocentric motion, language grounding, body tracking, daily activityopen
Aria Everyday Activities / AEA2024-02arXiv143 daily activity sequences across five indoor locations with trajectories, point clouds, gaze, speechEveryday AR perception, scene reconstruction, prompted segmentationopen
SANPO2023-09WACV 2025Scene understanding, accessibility, and human navigation data for egocentric navigation and spatial assistanceNavigation and assistive scene understandingwatch
Project Aria Datasets2023-08arXivOfficial portal for Aria-based datasets and toolingAR glasses, wearable sensing, scene reconstruction, gaze, SLAMopen
Aria Digital Twin / ADT2023-06arXiv200 real-world activity sequences, raw Aria streams, 6DoF poses, object poses, depth, segmentation, synthetic renderingsEgocentric 3D machine perception and digital-twin evaluationopen

Benchmarks and Derived Annotations

Evaluation suites and label sets built on top of the raw datasets above.

BenchmarkReleasedVenueBuilt onTasksStatus
SLT 2026 SmartGlasses Challenge2026-08IEEE SLT 2026 Challenge106.98 hours and 714 four-channel smart-glasses audio sessions released to registered participantsTimestamped speaker-attributed ASR and spoken-language understanding for dyadic dialogue and multi-party meetingsrequest
HandEdit2026-08arXiv300K+ clips and 200M+ editing instances from EgoDex, ARCTIC, OakInk2, HOI4D, and HO-CapURDF-conditioned human-to-robot hand/arm editing across 26 embodiments, with public data and evaluation codeopen
H2R-Bench2026-08arXivEgocentric human manipulation demonstrations paired with target robot embodimentsFive-dimension evaluation of 11 video generators across six manipulation families and two robot embodimentswatch
GST-Bench2026-08arXiv6,790 minutes of synthetic video with human-verified questionsGlobal spatial VQA from egocentric streams, unseen viewpoints, and top-down sceneswatch
EgoAfford2026-08arXiv15.5K human-verified generated images plus 102 real egocentric images across 26 tasksJoint next-step planning and role-specific affordance segmentationopen
HumanCLAW-Bench2026-07arXiv1,218 egocentric find-navigate-interact episodes across 41 indoor scenesVLM embodied decision-making, navigation, interaction, and body self-awarenesswatch
EgoSafe-Bench2026-07arXiv3,000 first-person mobile clips paired with hierarchical QA chains12,000 evaluations of evidence anchoring, blind spots, intent, causality, and safety reasoningwatch
VIABench2026-07arXivFirst-person videos recorded or shared by blind and visually impaired peopleProactive navigation reminders, assistive VideoQA, and vision-guided interaction in streaming and offline settingswatch
EgoProceVQA2026-07arXiv1,272 clips from CaptainCook4D, EPIC-Tent, Assembly101, and EgoOops3,600 QA pairs over 31 procedures and six key-step question types, plus EgoProceGen and EgoProceAgentwatch
EgoPolice2026-07arXivPublic police body-worn camera footageSecond-by-second high-stakes action labels, action classification, and multiple-choice QA over egocentric police-civilian interactionswatch
S-EMBER2026-07arXivRay-Ban Meta smart-glasses video3,141 videos, 388 hours, and 9,448 QA pairs for streaming episodic-memory retrieval with temporal evidencewatch
R3D-Bench / R3D2026-07arXivAria Digital Twin3,033 quantitative 3D spatial-reasoning questions over 57 posed egocentric RGB-D video sequences plus a tool-calling baselinewatch
EgoInertia-MI2026-07arXivSynchronized egocentric video and wearable IMUMotor-impairment assessment across 19 simulated upper- and lower-body activities and severity levelswatch
LongEgoRefer2026-07ECCV 2026Ego4D1,498 long-form video referring expressions with average 45-minute videos, sparse object occurrences, and spatio-temporal groundingopen
SG-Ego / GLEN2026-07arXivEgo4DSpatio-temporal scene-graph annotations, graph-text/action alignment, and activity-driven graph-edit forecastingwatch
NormAct2026-06arXivEgocentric embodied-planning scenesHidden social-norm compliance, goal achievement, and task-success evaluation for embodied plannerswatch
BinaryTracking / GangnamLoop2026-06arXivLong egocentric robot routesSpatial QA, metric coordinate retrieval, and navigation with an open BinTrack implementationopen
UMI-Bench 1.02026-06arXivUMI-style tabletop manipulationLocal-first real-robot benchmark protocol for UMI policy deployment and task-factor analysiswatch
EgoSAT2026-06arXivStreaming first-person interaction clipsEgocentric long-horizon interaction QA and decision-level alignment in real-time video streamswatch
NetraLink Assistive AI2026-06arXivHead-mounted GoPro assistive scenariosObject recognition, scene-text QA, and multilingual visual reading for wearable MLLMswatch
VEGA / VEGA-Bench2026-06arXivIn-the-wild egocentric navigation video with monocular geometry250K scenes and about 5M navigation goals for obstacle-aware VLA evaluationwatch
UCS-Bench / DirectMe2026-06ICML 2026170+ hours of egocentric observations8.1K+ timestamped questions for continual spatial intelligence and streaming spatial memoryopen
StreamMemBench2026-06arXivEgoLife egocentric streamsInitial/follow-up task sequences testing evidence recall, feedback incorporation, and future assistancewatch
V-RAGBench / CARVE2026-06arXivLong egocentric video chunksDecoupled retrieval and generation evaluation for VideoRAG over evidence chunkswatch
VL-MemKnG / WalkieKnowledgeT+2026-06arXivLong egocentric navigation trajectoriesTemporally distributed spatial-memory QA with hybrid graph and segment retrievalwatch
OVO-S-Bench2026-06arXivContinuous egocentric streamsHierarchical spatial intelligence across perception, tracking, simulation, and allocentric mappingwatch
SpatialWorld2026-06arXivInteractive real-world spatial tasks under vision-only partial observabilityEgocentric evidence gathering and spatial-agent evaluationwatch
HumanoidArena2026-06arXivEgocentric humanoid-control tasksHierarchical whole-body learning from egocentric perceptionwatch
Plan, Watch, Recover / EgoProactive2026-06arXivEgoProactive plus Pro2Bench over five established benchmarksProactive procedural assistance, out-of-plan detection, and recovery guidancewatch
Wearable AI Dataset2026-05ECCV 2026 Workshop2,100 gated head-mounted videos split evenly across EgoConv, EgoLongQA, and EgoProactiveConversational QA, long-video QA, and streaming proactive-assistant evaluation with a bundled starter kitrequest
PCSR-Bench2026-05arXiv2,600 omnidirectional indoor images84,373 perspective-conditioned spatial-reasoning questions, including egocentric rotation and limited-FOV visibilitywatch
EgoBench2026-05arXivEgocentric video tasksInteractive multimodal tool-using agentswatch
Minerva-Ego2026-05arXivEgocentric videos with dense reasoning traces and object masksMulti-step egocentric visual reasoningopen
EgoPro-Bench2026-05arXivPersonalized egocentric video streamsProactive interaction, intent timing, and personalized assistancewatch
EgoBabyVLM2026-05arXivNaturalistic infant and adult egocentric videosDevelopmental cross-modal VLM evaluationwatch
Beyond Motion Primitives2026-05arXivEgo4DHead-mounted IMU benchmark for behavioral activity recognition on smart glasseswatch
Ego-METAS2026-05arXivEgoExo4D, CMU-MMAC, CaptainCook4DOnline multimodal, energy-aware temporal action segmentation across RGB, audio, gaze, IMU, and monochrome streamswatch
EgoProx2026-05CVPR 2026Egocentric 3D proximity QAIntention, exploration, exploitation, and chain-of-actions spatial reasoning for MLLMswatch
BARISTA2026-05arXiv185 coffee-preparation videosScene graphs, masks, tracks, boxes, hand-object interactions, activities, and process-step reasoningwatch
TAVIS2026-05arXivIsaacLab active-vision imitation tasksHeadcam vs fixed-cam evaluation, wrist/head active vision, and anticipatory-gaze metricwatch
MM-Conv2026-05arXiv6.7 hours of egocentric VR interactionReferential communication with speech, motion, gaze, and 3D sceneswatch
EgoExoMem2026-05arXivSynchronized ego-exo videosCross-view memory QAwatch
EgoMemReason2026-05COLM 2026Week-long EgoLife video; 500 public questions plus evaluation code and leaderboardEntity, event, and behavior memory reasoningopen
Personal Visual Context Learning / Personal-VCL-Bench2026-05arXivContinuous smart-glasses streamsPrompt-time wearer-specific visual-context learning for personalized LMM assistantswatch
GazeMind / CogLoad-Bench2026-05arXivSmart-glasses gaze and cognitive-load annotationsGaze-guided LLM agent evaluation for personalized cognitive-load assessmentwatch
VIGIL2026-05arXivEgocentric RGB embodied-agent episodesTerminal-commitment scoring that separates world completion from success reportingwatch
Ego2World2026-05arXivHD-EPICExecutable symbolic worlds from egocentric cooking video for belief-state planningwatch
EgoLink 20262026-04ACM MM 2026E3 plus public challenge videos and labelsEgocentric social reasoning, emotion/intent/causality QA, and interactive tool-using agentsopen
SpaMEM2026-04arXivProcedural embodied environmentsDynamic spatial-memory updates from action-conditioned egocentric observationswatch
WhissleAI Egocentric Activity Sample2026-04Hugging Face19 public first-person clipsEgo4D-style narrations, NLQ, moment annotations, FHO actions, metadata, and taxonomy for prototypingopen
EgoPoint-Bench2026-04ACL 2026Simulated and real egocentric pointing samples11K+ QA items for referential reasoning and pointing-grounded object disambiguationwatch
PIE-V2026-04arXivEgo-Exo4D-style procedural scenariosMistake-aware procedural-video evaluation with recovery correctionswatch
ReFocus / EM-QnF2026-04CVPR 2026Egocentric episodic-memory NLQ with user feedbackInteractive feedback refinement for ambiguous memory querieswatch
EgoEsportsQA2026-04arXivFirst-person esports videoFast virtual first-person perception and reasoningwatch
Audio Hallucination in Egocentric Video2026-04ICASSP 2026Egocentric audio-visual video300 videos / 1,000 sound-focused questions probing audio hallucination in AV-LLMswatch
EXPLORE-Bench2026-03arXivReal first-person videosEgocentric long-horizon scene-state prediction and reasoningwatch
LifeEval2026-03arXivContinuous first-person streamsReal-time task-oriented human-AI collaboration in daily lifewatch
SAVA-X2026-03CVPR 2026EgoMe and asynchronous ego/exo video pairsEgo-to-exo imitation-error detectionwatch
HOI-Synth2026-03arXivVISOR, EgoHOS, and ENIGMA-51 with synthetic hand-object interaction annotationsSynthetic HOI detection and benchmarking under synthetic augmentationwatch
MA-EgoQA2026-03arXivMulti-agent egocentric streamsSocial, task coordination, theory-of-mind, temporal, environment QAopen
Ropedia Xperience-10M Task Suite2026-03Hugging FaceXperience-10M Sample12 embodied-AI task contracts, sample baselines, and evaluation protocolopen
Kriya-Egocentric-100K2026-03Hugging FaceEgocentric-100K preview subsetAction100M-style hierarchical temporal action trees and LLM-generated action descriptionsopen
EgoAVU2026-02CVPR 2026Egocentric audio-visual narrationsEgoAVU-Instruct (3M QAs) and EgoAVU-Bench (3K QAs) for audio-visual understanding (CVPR 2026 highlight)open
SAW-Bench2026-02ICML 2026Ray-Ban Meta smart-glasses videoObserver-centric situated awareness and physically grounded spatial reasoningopen
Ego4OOD2026-01arXivEgocentric video action recognitionCovariate-shift benchmark for egocentric domain generalizationwatch
Sanpo-D2026-01arXivSanpo egocentric navigation videoSpatial-conditioned reasoning over long first-person videos with fine-grained spatial re-annotationwatch
Egocentric-100K Evaluation2025-12Hugging FaceEgocentric-100K, Ego4D, EPIC-KITCHENS-10030K-frame comparison of hand visibility and active manipulation densityopen
Know-Show2025-12arXivCharades, Action Genome, Ego4DSpatio-temporal grounded reasoning with joint answer and evidence localizationwatch
Egocentric-10K Evaluation2025-11Hugging FaceEgocentric-10K, Ego4D, EPIC-KITCHENS-10030K-frame comparison of hand visibility and active manipulation density with published labeling promptsopen
CS3 Collision Sound Source Segmentation2025-11arXivEPIC-CS3 and Ego4D-CS3Audio-conditioned segmentation of objects responsible for collision sounds in egocentric videowatch
EgoExo-Con2025-10arXivSynchronized ego-exo videosView-invariant temporal verification and grounding; introduces View-GRPOwatch
LaMAria2025-09ICCV 2025Glasses-like wearable captureCity-scale egocentric visual-inertial SLAM with centimeter ground truthopen
EgoIllusion2025-08EMNLP 2025Egocentric videoHallucination benchmark probing fabricated objects, actions, and sounds in multimodal modelswatch
EgoCross2025-08AAAI 2026Cross-domain ego clips; the EgoVis 2026 challenge drew 1,500+ submissions from 130+ participantsSurgery, industry, extreme sports, animal-perspective QA, and two Codabench trackswatch
EgoExoBench2025-07arXivPublic ego-exo video datasets7,300+ QA pairs over semantic alignment, viewpoint association, and temporal reasoningwatch
EASG-Bench2025-06ICCV 2025 WorkshopEgocentric action scene graphsRelation and temporal QAopen
HD-EPIC VQA Challenge2025-02CVPR 2025HD-EPICRecipe, ingredient, nutrition, fine-grained action, 3D perception, object motion, gazebenchmark
EFM3D2024-06arXivProject Aria egocentric 3D data3D object detection and surface regression for egocentric foundation modelswatch
Ego4D-OSCA2024-05arXivEgo4DObject-state-change anticipation annotations and visual-language baseline for procedural videowatch
EgoHOIBench / EgoNCE++2024-05ICLR 2025Multiple egocentric HOI sourcesOpen-vocabulary hand-object interaction understandingopen
EgoPlan-Bench2023-12IJCVEgocentric planning tasksMultimodal LLM planning over human-level egocentric scenariosopen
Ego-Exo4D Benchmarks2023-11project pageEgo-Exo4DFine-grained activity, proficiency, cross-view translation, 3D pose, object correspondencebenchmark
Egocentric Pedestrian Trajectory Benchmark2023-10ICRA 2024Egocentric pedestrian videoTrajectory prediction with scale- and motion-aware evaluationwatch
EPIC-Aff / Multi-label Affordance Mapping2023-09ICCV 2023EPIC-KITCHENSMetric, spatial, multi-label affordance annotations for interaction-grounded segmentation and navigation mapsbenchmark
RefEgo2023-08ICCV 2023Ego4D12K+ clips / 41 hours for first-person referring-expression comprehension and referred-object trackingopen
EgoSchema2023-08NeurIPS 2023Ego4DVery-long-form multiple-choice video QAopen
EgoAdapt 20232023-07arXivEgo4DReal-world online adaptation benchmark over 50 user streams and 2,740 action labelsbenchmark
EPIC-Fields2023-06NeurIPS 2023EPIC-KITCHENS3D fields and scene-level spatial reasoning over kitchen videoopen
MMG-Ego4D2023-05CVPR 2023Ego4DMissing-modality and cross-modal zero-shot generalization benchmark over video, audio, and IMUbenchmark
Fine-Grained Affordance Annotation2023-02WACV 2023Egocentric HOI videosFine-grained affordance labels for hand-object interactionwatch
EPIC-Sounds2023-02ICASSP 2023EPIC-KITCHENSAudio event recognition in egocentric kitchen videoopen
VISOR2022-09NeurIPS 2022EPIC-KITCHENSManual and dense masks, hand/object segmentation, active object relationsopen
EgoClip / EgoMCQ2022-06NeurIPS 2022Ego4D3.8M clip-text pairs and MCQ development benchmark for egocentric VLPopen
AssistSR2021-11EMNLP 2021Instructional daily-item video segmentsTask-oriented question-driven video segment retrieval for personal assistantswatch
Ego4D Benchmarks2021-10project pageEgo4DNatural Language Query, Moment Query, episodic memory, state change, long-term anticipation, social/audio, hand-objectbenchmark
TREK-1502021-08project pageEPIC-KITCHENSEgocentric single-object trackingopen
EPIC-KITCHENS Challenges2018-04project pageEPIC-KITCHENS / EPIC-KITCHENS-100Recognition, detection, anticipation, retrieval, domain adaptationbenchmark

Models, Tools, and Baselines

Open models, baselines, and loaders you can build on directly.

Egocentric Foundation Models and Assistants

ResourceReleasedVenueWhat it contributesLink
Vision-Language Models for Egocentric Video2026-08arXivCritical survey connecting egocentric VLMs across hand-object interaction, temporal and graph reasoning, efficient long-video processing, wearable assistance, and embodied AIPaper
EgoGazeLite2026-08ECCV 2026 WearableAI WorkshopSoftware-only gaze prediction for token-efficient wearable MLLM input with 15.7M parameters and 21.6 ms/frame end-to-end latencyPaper
Long-Tail Embodied Urban Navigation from In-the-Wild Videos2026-08arXivScalable annotation of web egocentric video with metric trajectories and navigation semantics for VLA training and rare-failure analysisPaper
Hand-Object Interaction in the Age of Large Foundation Models2026-07arXivSystematic survey of foundation-model priors across six HOI tasks and eight geometric, semantic, and visual prior types, including embodied transferPaper
Ego Scene Augmentation (ESA)2026-07arXivExplicit ego-element graph that fuses visual-foundation-model signals for spatially structured scene understanding and stronger indoor/outdoor EgoTextVQAGitHub
Worldscape-MoE2026-07arXivMoE diffusion world model sharing dynamics across camera trajectories, robot actions, and egocentric hand-joint control; code and weights are announced but not yet releasedProject
UNIEGO2026-06arXivUnified egocentric encoder distilled from nine teachers spanning ego-exo views, RGB, depth, skeleton, and foundation-model representationsPaper
ActiveMimic2026-06arXivEgocentric human-video pretraining with active-perception signals for manipulation and VLA transferPaper
Wh02026-06arXivWorld-model framework that synthesizes realistic hand-object manipulation sequences to bootstrap scalable egocentric data and policy learningPaper
VLESA2026-06arXivGoal-conditioned safety-assistance evaluation and baseline for monitoring human activities from egocentric videoGitHub
Continual Child-View Learning2026-06arXivChronological multimodal learning from a child's egocentric video and speech streamPaper
Objects Before Words2026-06arXivObject-first language grounding from child-view egocentric videoPaper
Watch Remember Reason2026-06arXivHuman-view long-video understanding framework for MLLM watching, memory, and reasoningPaper
World Action Models2026-05arXivSurvey and taxonomy linking VLA, world models, portable human demonstrations, simulation, and internet-scale egocentric videoPaper
Pro2Assist2026-05arXivContinuous step-aware proactive assistance framework with multimodal egocentric perception and AR-glasses evaluationPaper
T-REN2026-04arXivText-aligned region-token encoder that reduces long-video token counts and improves Ego4D video object localizationGitHub
V-JEPA 2.12026-03arXivDense self-supervised video representation model with strong Ego4D STA, EPIC-KITCHENS anticipation, and robot-grasping transfer resultsPaper
EgoViT Object SSL2026-03CVPR 2026Self-supervised object representation learning from continuous, uncurated first-person videoPaper
RynnBrain2026-02arXivOpen embodied foundation model family with variants for egocentric understanding, localization, physical reasoning, planning, navigation, and VLAPaper
Central Vision SSL2026-02arXivEgo4D-derived gaze-centered crops and temporal-slowness SSL for object representation learning from human-like visual experienceGitHub
PhysBrain2025-12arXivUses human egocentric data to bridge vision-language models toward physical intelligence and embodied controlPaper
EgoM2P2025-06ICCV 2025Egocentric multimodal multitask pretraining over RGB, depth, gaze, and camera posePaper
Exo2Ego / Ego-ExoClip2025-03AAAI 2026Transfers exocentric MLLM knowledge into egocentric video understanding with 1.1M synchronized ego-exo clip-text pairs and EgoIT instruction tuningPaper
GazeLLM2025-03Augmented Humans 2025Uses gaze-focused first-person regions to reduce MLLM video processing while preserving task comprehensionPaper
EgoHOD2025-03ICLR 2025Fine-grained hand-object-dynamics pretraining for egocentric representations (ICLR 2025)GitHub
EgoSpeak2025-02NAACL 2025 FindingsReal-time speech-initiation prediction for egocentric conversational agents over first-person RGB streaming videoPaper
PRVQL2025-02ICCV 2025Progressive knowledge-guided refinement for egocentric visual query localization under appearance change and clutterPaper
Human Gaze Object-Centered Learning2025-01arXivGaze-inspired central visual amplification for object-centered representation learning from Ego4D-scale egocentric experiencePaper
GEM Ego-Vision World Model2024-12CVPR 2025Generalizable ego-vision world model trained on 4,000+ hours for RGB/depth generation and ego-trajectory controlPaper
Vinci2024-12IMWUT 2025Real-time, always-on egocentric assistant with historical-context QA, planning, and visual demosGitHub
Ego-VPA2024-07WACV 2025Parameter-efficient adaptation for egocentric video-language foundation models with sparse basis promptsPaper
EgoVideo2024-06CVPR 2024Egocentric video foundation model with slow-fast adaptation; multi-track Ego4D / EPIC challenge winnerGitHub
AlanaVLM2024-06EMNLP 2024Embodied-AI foundation model for egocentric video understandingPaper
Missing Modality Token / MMT2024-01CVPRW 2025Missing-modality modeling for multimodal egocentric action recognition and moment localization on Ego4D, EPIC-KITCHENS, and EPIC-SoundsPaper
EgoDistill2023-01tech reportDistills heavy egocentric clip features into efficient models using head-motion signalsProject
EgoT2 / Egocentric Video Task Translation2022-12CVPR 2023 HighlightTask-translation framework that transfers supervision between egocentric video tasksProject

Video-Language and Long-Video Models

ResourceReleasedVenueWhat it contributesLink
MERIT2026-08ECCV 2026 OralSimple multi-key episodic memory and query-time neighbor filtering for high-recall ultra-long video retrieval; reports 71.2 EgoLifeQA accuracy with a GPT-5 backboneProject
SCOUT2026-08ACM MM 2026Recovery-aware tool-thought agent that self-checks retrieval observations, switches temporal regions, and trains with uncertainty-prioritized policy optimizationPaper
R4DSG2026-08ACM MM 2026Relative 4D scene-graph memory for persistent objects, anchor-relative state changes, and object-centric questions in long EgoLife videoProject
EgoCITE2026-08arXivContext-augmented multi-view memory indices plus question-conditioned temporal retrieval over EgoLifeQA, EgoMem, and EgoR1-BenchPaper
SmartRes2026-08arXivDynamic pixel-space resolution routing cuts visual tokens by up to 67% on Ego4D/EgoIntention grounding while retaining 86.4% of full-resolution performancePaper
EgoPlay2026-07SIGGRAPH Asia 2026Event-triggered egocentric video editing trained on 106K clip-prompt pairs, jointly learning trigger recognition, temporal restraint, and post-event editingProject
VideoTreeSearch2026-07arXivSelf-correcting temporal-tree agent for grounded long-video QA, with MIT-licensed inference/SFT/RL code, an 8B checkpoint, and released Haystack-Ego4D trajectories/evaluation dataGitHub
OmniView-Space2026-07arXivTool-guided egocentric spatial reasoning with query-aligned visual cognitive maps, textual spatial graphs, and ego-frame rewardsPaper
CGGS2026-07IEEE TIP 2026Text-to-3D egocentric scene generation with consistency-augmented 2D priors, point-track/depth layout decoration, and geometric Gaussian refinementGitHub
Egocentric Scene Graphs / EgoSG2026-06arXivConverts long-form egocentric videos into temporally grounded scene graphs for compact HD-EPIC VQA reasoningPaper
Low-Latency Doubly-Correct VLMs2026-06arXivRationale-informed pruning for efficient egocentric VLMs while preserving both answer correctness and evidence groundingPaper
DR-MV3D2026-06arXivDense-reward multi-view 3D VQA pipeline with global-map construction, view planning, and egocentric groundingPaper
Temporal Action Graphs for Ego VLMs2026-06arXivConverts egocentric videos into narratives and temporal action graphs for in-context action recognition with open-weight VLMsPaper
ReRe Cross-View Revisiting2026-06ICML 2026Training-free spatial reasoning that revisits egocentric conclusions through synthesized complementary novel-view videosProject
CASTLE KG Retrieval2026-06CVPR 2026 EgoVisAgentic long-context video understanding with video knowledge graphs and hierarchical retrieval for the CASTLE challengePaper
AnchorWorld2026-06arXivEmbodied egocentric world simulation with view-based evolution customizationPaper
Understanding-Enhanced Ego Mistake Detection2026-06arXivSmall/large-model collaboration for detecting incorrect procedural actions in egocentric video, with a Qwen3-VL Embedding reasoning branchPaper
FlexLAM2026-06arXivVariable-length latent actions for action-free video interfaces, improving scarce-label alignment and Ego4D transition reconstructionPaper
VisualClaw2026-06arXivReal-time personalized multimodal agent with streaming-frame filtering, skill evolution, and VisualClawArena workspace evaluationPaper
FlowNar2026-05ICML 2026Scalable streaming narration with bounded visual memory, evaluated on long-form Ego4D, Ego-Exo4D, and EPIC-KITCHENS-100-style narrationGitHub
E3C Video Generation2026-05arXivControllable egocentric video diffusion with 3D environmental memory and ego-exo human-pose controlProject
CASTLE2026 Team WDL2026-05CVPR 2026 EgoVisEvidence-aware multimodal reasoning pipeline for long-form CASTLE egocentric QAPaper
CuriosAI CASTLE2026-05CVPR 2026 EgoVisSearch-verify-answer CASTLE challenge pipeline using timelines, transcripts, and VLM captionsPaper
MARS CASTLE2026-05CVPR 2026 EgoVisMultimodal agentic reasoning with source selection for CASTLE challenge QAPaper
OSGNet + MLLM Reranking2026-05CVPR 2026 EgoVisChampion Ego4D Episodic Memory Challenge solution for NLQ and GoalStep using MLLM reranking over OSGNet candidatesGitHub
OmniEgo-R22026-05CVPR 2026 EgoVisRouted reasoning framework for EgoCross; second place in both Source-Limited and Open-Source tracksGitHub
Being-H0.72026-05arXivLatent world-action model from egocentric videos for future-aware reasoning and VLA policy learningPaper
Reflective Dialogue EgoCross2026-05CVPR 2026 EgoVisInference-time Teacher/Solver reflective dialogue for EgoCross support-set adaptation without fine-tuningPaper
EgoCross Domain-Wise Inference2026-05CVPR 2026 EgoVisNearly training-free source-limited inference strategy for EgoCross domain shiftPaper
HD-EPIC Semantic-Visual Evidence2026-05CVPR 2026 EgoVisHD-EPIC VQA challenge solution separating semantic and visual evidencePaper
HiERO-StepG2026-05CVPR 2026 EgoVisHierarchical activity-understanding solution for the Ego4D Step Grounding ChallengePaper
TempRet2026-05CVPR 2026 EgoVisTemporal enhancement and reranking for EPIC-KITCHENS-100 multi-instance retrievalPaper
Trajectory-Conditioned Egocentric Prediction2026-05arXivFuture ego-view prediction conditioned on camera trajectory to disambiguate action outcomesPaper
EventPrune2026-05arXivEvent-camera-guided visual-token pruning for efficient first-person dynamic spatial reasoningPaper
SpatioRoute2026-05arXivQuestion-aware prompt routing for zero-shot egocentric spatial QA without fine-tuning or 3D inputPaper
MAGIC-Video2026-05arXivMultimodal memory graph plus narrative-chain retrieval for ultra-long egocentric video reasoningGitHub
SpaceMind++2026-05arXivBuilds allocentric cognitive maps from RGB video to support spatially consistent reasoning over fragmented egocentric observationsPaper
EgoSim2026-04ECCV 2026Closed-loop egocentric world simulator with 3D grounding and dynamic state updates; inference code and weights are public while training code and a license remain pendingProject
World2VLM2026-04arXivDistills world-model imagination into VLMs for dynamic spatial reasoning under egocentric motionPaper
EgoMotion2026-04arXivHierarchical reasoning + diffusion for egocentric vision-language motion generationPaper
Ego-InBetween2026-04CVPR 2026Generates object state transitions in ego-centric videos from action instructionsPaper
EgoTSR2026-04arXivCurriculum-based task-oriented egocentric spatiotemporal reasoning with a reported 46M-sample EgoTSR-Data training setPaper
Syn2Seq Exo-to-Ego2026-04arXivSequential exo-to-ego video generation from synchronized third-person views and camera posesPaper
UniversalVTG2026-04arXivLightweight cross-dataset foundation model for video temporal groundingPaper
V-Nutri2026-04CVPR 2026 MetaFood WorkshopDish-level nutrition estimation from egocentric cooking videosPaper
LOME2026-03arXivAction-conditioned egocentric world model generating photorealistic human-object interactions from image, text, and per-frame actionsPaper
Temporal-Aware Ego VLM2026-03arXivTraining scheme that incentivizes temporal awareness in egocentric video-understanding modelsPaper
EgoForge2026-03arXivGoal-directed egocentric world simulator that rolls out first-person video from a single image and a high-level instructionPaper
EgoReasoner2026-03arXivTask-adaptive structured thinking for egocentric 4D spatial and object reasoningPaper
Gaze-Regularized VLMs2026-03arXivGaze-conditioned VLM training for ego-centric behavior understandingPaper
Recurrent Reasoning VLM2026-03CVPR 2026Recurrent VLM reasoning for long-horizon embodied task-progress estimationPaper
Ropedia Xperience-10M Task Baselines2026-03Hugging FacePublic Xperience-10M sample task definitions, baselines, metrics, and scale-up evaluation notesHugging Face
EgoMAS2026-03arXivShared-memory baseline for multi-agent egocentric video QAProject
DreamDojo2026-02ICML 2026Generalist robot world model pretrained on 44K hours of egocentric human video with continuous latent actions, distilled to real timeGitHub
Hand2World2026-02arXivAutoregressive egocentric world model generating first-person interaction video from free-space hand gesturesPaper
Edge Episodic Memory QA2026-02VISAPP 2026On-device dual-thread MLLM for real-time egocentric episodic-memory QA (QAEgo4D-Closed) on the edgePaper
EgoGraph2026-02arXivTraining-free temporal knowledge graph for ultra-long egocentric video understandingPaper
EGAgent2026-01arXivEntity-scene-graph agent for very-long egocentric video understanding over all-day wearable streamsGitHub
Walk through Paintings2026-01arXivEgocentric world models from internet priors that generate first-person scene walkthroughsPaper
Event-VStream2026-01arXivEvent-driven long-video stream understanding for real-time video-language systemsPaper
EgoHandICL2026-01ICLR 2026First-person hand-object in-context learning from multimodal egocentric context and synthetic exemplarsPaper
HD-EPIC VQA T-CoT2026-01CVPR 2025 EgoVisHD-EPIC VQA solution with temporal chain-of-thought prompting and Qwen2.5-VL adaptationPaper
Robust Egocentric Visual Attention2026-01arXivLanguage-guided scene-context model for egocentric visual attention predictionPaper
The N-Body Problem2025-12ECCV 2026 WorkshopPredicts feasible multi-person parallel execution plans from a single EPIC-KITCHENS or HD-EPIC egocentric videoProject
EgoVITA2025-11ECCV 2026Plan-then-verify egocentric video reasoning with visual-grounding and cross-perspective consistency rewardsPaper
EgoControl2025-11arXivPose-controllable egocentric video diffusion conditioned on sequences of 3D full-body posesPaper
EAGLE VQL2025-11AAAI 2026Episodic appearance- and geometry-aware memory for unified 2D/3D visual query localization in egocentric videoPaper
TimeSearch-R2025-11arXivReinforcement-learned temporal search with self-verification for Haystack-Ego4D and long-video understandingGitHub
DMC3 Ego VideoQA2025-10ACM MM 2025Counterfactual contrastive construction for egocentric VideoQA across event descriptions and hand-object interaction cuesPaper
EgoThinker2025-10NeurIPS 2025Egocentric reasoning model with spatio-temporal chain-of-thought and RL fine-tuningPaper
HieraMamba2025-10CVPR 2026Hierarchical Anchor-Mamba pooling for long-video temporal grounding, preserving fine temporal structure on Ego4D NLQ-style queriesProject
EgoPrompt2025-08ACM MM 2025Prompt-learning framework for egocentric action recognition that models verb-noun semantic and contextual relationshipsPaper
EgoTwin2025-08arXivJoint first-person video and human-motion generation with head-centric motion alignmentPaper
Ego-PM2025-08arXivEgocentric predictive model that jointly forecasts future actions and video frames conditioned on hand trajectoriesPaper
GazeNLQ2025-06CVPR 2025 Ego4D ChallengeGaze-augmented Ego4D Natural Language Queries localization from estimated wearer attentionPaper
EgoWorld2025-06ICLR 2026Reconstructs egocentric views from exocentric point clouds, 3D hand poses, and text for cross-view world generationProject
EgoVLM2025-06arXivGRPO-style policy optimization over Qwen2.5-VL-3B for egocentric video reasoningPaper
Ego-R12025-06arXivChain-of-tool-thought framework and training data for ultra-long egocentric video QAPaper
HiERO2025-05ICCV 2025Weakly supervised hierarchical activity-thread representations for EgoMCQ, EgoNLQ, and procedure learningPaper
OSGNet2025-05CVPR 2025Object-shot enhanced grounding for egocentric temporal grounding / moment queries (CVPR 2025)GitHub
EgoExo-Gen2025-04ICLR 2025Predicts future ego-centric video from exocentric video, first ego frame, text instructions, and hand-object masksPaper
EgoDTM2025-03arXivDepth- and text-aware egocentric VLP for 3D-aware representationsGitHub
EgoLife / EgoButler / EgoGPT / EgoRAG2025-03arXivOmni-modal egocentric assistant system with retrieval over long life recordingsPaper
EgoAgent2025-02ICCV 2025Joint predictive agent model for perception, future-state prediction, and action in egocentric worldsPaper
Embodied VideoAgent2024-12ICCV 2025Builds persistent 3D scene memory from egocentric video plus depth/pose sensors for dynamic scene reasoning and planningProject
ESOM / OVQ2D2024-11WACV 2026Online visual query localization with compact egocentric streaming object memoryPaper
MM-Ego2024-10ICLR 2025Egocentric multimodal LLM with Memory Pointer Prompting and a 7M-sample egocentric QA data engine (EgoMemoria benchmark)Paper
AMEGO2024-09ECCV 2024Active-memory representation from a single long egocentric video for fast multi-query answering (ECCV 2024)Project
EgoNCE++2024-05ICLR 2025Open-vocabulary HOI benchmark and asymmetric contrastive objectiveGitHub
SoundingActions2024-04CVPR 2024Self-supervised audio-language-vision embedding that learns how actions sound from narrated egocentric videosProject
Video ReCap2024-02CVPR 2024Recursive captioning for hour-long videos plus Ego4D-HCap long-range summariesProject
EgoInstructor2024-01CVPR 2024Retrieval-augmented egocentric captioning via cross-view retrieval of third-person clips (CVPR 2024)Project
GroundVQA2023-12CVPR 2024Unified temporal grounding and open/closed QA for long egocentric videosGitHub
LifelongMemory2023-12arXivLLM-based long-form egocentric memory system for answering queries over a camera wearer's pastPaper
LEGO2023-12ECCV 2024Visual-instruction-tuned egocentric action-frame generationPaper
Action Scene Graphs2023-12CVPR 2024Temporally evolving action scene graphs for long-form egocentric video understandingPaper
LEAP2023-11arXivLLM generation of egocentric action programs with sub-actions, preconditions, postconditions, and object referencesPaper
PALM2023-11ECCV 2024Predicts future actions through language models for long-term action anticipationPaper
Exo2EgoDVC2023-11WACV 2025Dense video captioning for egocentric procedural activities using web instructional videosProject
CliMer Temporal Grounding2023-10BMVC 2023Learns temporal sentence grounding from narrated Ego4D/EPIC videos using clip merging and rough narration timestampsGitHub
Encode-Store-Retrieve2023-08ISMAR 2024Language-encoded egocentric memory pipeline for storing and retrieving human-perception eventsPaper
EgoVLPv22023-07ICCV 2023Egocentric video-language pretraining with fusion in the backboneProject
VQLoC2023-06NeurIPS 2023Single-stage visual query localization with joint query-frame and frame-frame correspondences for long egocentric videoProject
EgoCOL2023-06CVPR 2023 Ego4D ChallengeCamera-pose estimation for Ego4D open-world 3D object localization using sparse video/scan pose reconstructionGitHub
SpotEM2023-06ICML 2023Efficient episodic-memory search that keeps most NLQ accuracy while computing only a small fraction of clip featuresProject
GroundNLQ2023-06CVPR 2023Two-stage multi-scale grounding for Ego4D Natural Language Queries; CVPR 2023 NLQ championGitHub
MINOTAUR2023-02arXivUnified video grounding model for multimodal queries over Ego4D-style grounding tasksPaper
NaQ2023-01CVPR 2023Uses narrations as implicit queries to supervise episodic-memory retrieval in long egocentric videoPaper
LaViLa2022-12CVPR 2023Learns video-language representations from Ego4D narrations and LLM-generated narrationsPaper
CocoFormer VQL2022-11Ego4D Challenge 2022Proposal-set transformer for Ego4D visual query localization, improving VQ2D/VQ3D object localizationGitHub
Negative Frames Matter2022-08Ego4D Challenge 2022Trains Ego4D VQ2D with noisy negative/background frames and a more efficient object-query training loopGitHub
EgoEnv2022-07NeurIPS 2023 OralHuman-centric environment representations linking egocentric video to the wearer's local surroundingsProject
EgoVLP / EgoNCE2022-06NeurIPS 2022EgoClip, EgoNCE objective, EgoMCQ, and transfer to Ego4D/EPIC/Charades-Ego tasksGitHub

Action, Tracking, Pose, and HOI Baselines

ResourceReleasedVenueWhat it contributesLink
DreamHand2026-08arXivDeterministic diffusion geometry encoder for continuous metric bimanual trajectories under occlusion and out-of-sight gaps, with large gains on ARCTIC and HOT3DProject
G3Ego2026-08ECCV 2026 CONTEXTUS WorkshopGaze-guided pruning of action-scene graphs for efficient and interpretable recognition and anticipation on EGTEA Gaze+ and MECCANOPaper
EgoHRV2026-08ECCV 2026Heart-rate and HRV estimation from headset gaze cameras; reported HRV features improve EgoExo4D proficiency classification by 17.8%Paper
EgoTac2026-08arXivPredicts continuous force and contact from egocentric video after training on more than 5.7M image-tactile pairsPaper
EgoTrack3D2026-08arXivDynamic 3D object tracking from egocentric RGB through global-frame mask lifting, motion scoring, and voxel association; reports 11% higher PCL on Aria Digital TwinPaper
EgoHieraLoc2026-08arXivUnified 2D/3D visual-query localization with segmentation-guided representations and geometry-semantic confidence across viewpointsPaper
EgoPHI2026-08ECCV 2026Dense hand/object contact maps and 3D force distributions from one ego RGB image and object geometry, with simulation supervision and an eight-participant physical evaluationPaper
Clinical Freezing-of-Gait with Egocentric Vision2026-08ECCV 2026 WorkshopAt-home comparison of ego-video and IMU representations using synchronized recordings and expert labels from 13 people with Parkinson's diseasePaper
HiResNets2026-08arXivNative Full-HD foveal residual streams preserve fine egocentric objects without quadratic feature-grid growthPaper
Trustworthy Visual Predicates2026-06arXivReliability framework for contact, grasp, release, support, and other manipulation predicates under visual degradation, evaluated on VISOR/EPIC-KITCHENS, H2O, and ARCTICPaper
FactCheck LTA2026-06arXivFeasibility-aware long-term action anticipation with a multi-agent Observe-Plan-Verify loop over EPIC-KITCHENS-55 and EGTEA Gaze+Paper
TrAction2026-06arXivSparse 2.5D point-trajectory transformer for action recognition that reduces appearance/background shortcuts on EPIC-KITCHENS-100GitHub
HumanScale2026-06arXivShows filtered, labeled egocentric human video can outperform teleoperated real-robot data for embodied pretrainingPaper
Do as I Do2026-06arXivRetargets human hand-object interactions from in-the-wild monocular video into executable dexterous robot trajectoriesPaper
Motion-Focused Latent Action VLA2026-06IROS 2026Extracts motion-focused latent actions from unlabeled human EgoVideos for cross-embodiment VLA pretraining and adaptationPaper
EgoPriMo2026-06arXivEgocentric motion prior for interactive humanoid control from human demonstrationsPaper
EgoPhys2026-06arXivLearns deformable-object physics digital twins from egocentric RGB interaction video for robot planningProject
EDITH2026-06arXivStreams first-person view, gaze, and speech from smart glasses into hierarchical robot policies for natural HRIProject
Divide Deliberate Decide2026-06arXivLocal zero-shot multi-agent VLM framework for fine-grained egocentric action recognitionPaper
Ego-Nav Co-training2026-06arXivConverts egocentric walking videos into robot-action datasets and co-trains a VLA with robot demos for mobile navigationPaper
EgoGuide2026-06arXivSynchronized wrist/head egocentric demonstration collection with online guidance and a gated egocentric residual policyProject
Ego-Pi2026-06arXivVLA fine-tuning study across egocentric human and robot data using pi0.5 and dexterous five-finger embodimentsProject
HALOMI2026-06arXivExtends UMI-style egocentric collection for humanoid loco-manipulation with active perception from human demonstrationsPaper
EgoTactile2026-06ICML 2026 spotlightEgocentric video to full-hand grasp-pressure benchmark and diffusion/baseline models for everyday-object interactionsProject
EgoPressDiff2026-06ICASSP 2026Conditional video diffusion for UV-domain egocentric hand-pressure maps using pose, mesh, and depth conditioningProject
Hand Trajectory Fusion for Ego NLQ2026-06CVPR 2026 EgoVisHand-trajectory encoder and cross-attention fusion for Ego4D Natural Language Query groundingPaper
PROSE2026-06arXivTraining-free RGB-only egocentric scene registration using VLM-derived object-level 3D scene graphsProject
ACE-Ego-02026-06arXivUnifies egocentric human video with robot/sim data via reliability-aware weighting for VLA pretraining (RoboCasa GR1, RoboTwin 2.0)Paper
HumanNet2026-05arXiv1M-hour mixed first-/third-person human-centric corpus with interaction annotations and a 1K-hour egocentric subset used for VLA validationPaper
EgoAction2026-05CVPR 2026CVPR 2026 EPIC-KITCHENS action detection challenge pipelinePaper
EgoAdapt2026-05CVPR 2026CVPR 2026 HD-EPIC VQA challenge inference-time adaptation pipelinePaper
FROST-STA2026-05CVPR 2026 EgoVisFrozen dense-feature Ego4D short-term object-interaction anticipation submissionPaper
JFAA2026-05CVPR 2026 EgoVisJEPA-based EPIC-KITCHENS-100 action-anticipation challenge submissionPaper
Mamba Ego Action Recognition2026-05CVPR 2026 EgoVisCross-modal egocentric action recognition using RGB and hand-skeleton streams with MambaPaper
TAP-JEPA2026-05CVPR 2026 EgoVisEPIC-KITCHENS-100 action-anticipation runner-up using frozen V-JEPA 2.1 featuresPaper
VISTA2026-05CVPR 2026 EgoVisV-JEPA plus StillFast temporal anticipator for Ego4D STA at EgoVis 2026Paper
WristCompass2026-05arXivLearns ego-camera orientation from hand/camera kinematic coupling in manipulation videoPaper
Zero-Shot Ego Object ReID2026-05arXivSAM3-feature fusion for zero-shot object re-identification in egocentric kitchen videosPaper
EgoRelight2026-05arXivHMD-based egocentric human capture and illumination recovery for relightable avatarsPaper
StableHand2026-05arXivQuality-aware flow-matching baseline for world-space dual-hand motion estimation from egocentric videoProject
HumanEgo2026-05arXivZero-shot robot learning from minutes of human egocentric video; public code and 122 Aria recordings include raw VRS/MPS plus processed entity-level hand-object supervisionProject
EgoSPT / SPOT2026-05arXivSpatially prompted egocentric manipulation trajectory prediction from first-frame object/target groundingPaper
EggHand2026-05CVPR 2026 FindingsMultimodal egocentric hand-pose forecasting using video-language and VLA-style action decodingProject
Map-Mono-Ego2026-05ICIP 2026Map-grounded global human-pose estimation from monocular egocentric video and scanned environmentsPaper
EgoForce Hand Pose2026-05SIGGRAPH 2026Monocular egocentric 3D hand pose and shape reconstruction across fisheye, perspective, and wide-FOV camera modelsProject
EARL2026-05ICML 2026Analysis-guided RL framework for egocentric interaction reasoning and pixel grounding with coarse-to-fine parsingGitHub
EgoExo-WM2026-05arXivConverts exocentric video into egocentric world-model training data using body-pose priorsProject
MotionGRPO2026-05ICML 2026GRPO-based post-training for full-body 3D motion recovery from head-mounted device signalsPaper
ActiveGlasses2026-04arXivLearns robot manipulation from smart-glasses ego-centric human demonstrations and transfers active vision to a robot perception armPaper
GazeVLA2026-04arXivPretrains on egocentric gaze, intention, and action signals before robot fine-tuning for manipulationProject
WARPED2026-04arXivWrist-aligned rendering turns monocular egocentric human demonstrations into robot policy observationsPaper
DP-DeGauss2026-04ICASSP 2026Dynamic probabilistic Gaussian decomposition for egocentric 4D scene reconstruction of hands, objects, and backgroundPaper
Personal Point of View 3DGS2026-04CVPR 2026 EgoVisEvaluation of dynamic 3D Gaussian splatting for egocentric scene reconstructionPaper
VGGT-Segmentor2026-04arXivGeometry-enhanced segmentation across egocentric and exocentric viewsPaper
Gaze-SoM HOI Anticipation2026-04ICPR 2026Gaze and set-of-mark prompting in VLLMs for hand-object-interaction anticipation from egocentric videoPaper
EgoFlow2026-04CVPR 2026Gradient-guided flow matching for physically plausible 6DoF object-motion generation from egocentric videoPaper
EgoAdapt Speaker Detection2026-03arXivRobust egocentric Talking-to-Me speaker detection under missing modalities, head-orientation ambiguity, and background noisePaper
UniDex2026-03CVPR 2026Robot foundation suite for universal dexterous hand control learned from egocentric human videosPaper
PAWS2026-03arXivArticulation extraction from large-scale hand-object interactions in egocentric videoPaper
AG-EgoPose2026-03arXivAttention-guided egocentric 3D human-pose estimation from fisheye camera inputPaper
Static Scene Reconstruction from Dynamic Egocentric Videos2026-03arXivMask-aware 3D reconstruction pipeline for long-form dynamic egocentric videoPaper
EgoHOI World Model2026-03arXivPhysics-informed egocentric world model that synthesizes contact-consistent hand-object interactions from action signals alonePaper
STAformer++ Affordance-Aware Anticipation2026-02IEEE TPAMIIntegrates temporal attention, scene affordance memory, and interaction hotspots for short-term object-interaction anticipation on Ego4D and EPIC-KITCHENSPaper
Neck-Mounted Gaze (GLC)2026-02arXivTransformer gaze estimator for a shoulder-level neck-mounted camera with out-of-bound classification and multi-view co-learningPaper
EgoPush2026-02arXivEnd-to-end egocentric multi-object rearrangement for mobile robots from a single first-person camera, with zero-shot sim-to-realPaper
DeltaDorsal2026-01CHI 2026Dorsal hand-skin deformation features for robust egocentric hand pose under self-occlusion (~18% lower joint-angle error)GitHub
DexWM2025-12CVPR 2026Dexterous interaction world model from finger keypoints in 900+ hours of egocentric human and robot videoPaper
WholeBodyVLA2025-12ICLR 2026Unified latent VLA learning latent actions from action-free egocentric videos for whole-body humanoid loco-manipulationGitHub
GateFusion2025-12WACV 2026Hierarchical gated cross-modal fusion for active speaker detection, including unconstrained and egocentric ASD settings such as Ego4D-ASDPaper
EgoSpanLift2025-11NeurIPS 2025 spotlightLifts 2D egocentric gaze forecasting into 3D visual-span prediction (SLAM keypoints + 3D U-Net) over a 364.6K-sample benchmarkPaper
Uni-Hand2025-11T-PAMI 2026Universal egocentric hand motion forecasting for 2D/3D wrist and finger waypoints with code and dataGitHub
SFHand / EgoHaFL2025-11arXivStreaming language-guided 3D hand forecasting plus synchronized 3D hand-pose and language-instruction dataPaper
EgoCogNav2025-11arXivCognition-aware egocentric navigation forecasting trajectory and head motion from perceived path uncertainty (CEN dataset)Paper
In-N-On2025-11arXivScales egocentric manipulation with 1,000+ hours in-the-wild human video (PHSD) to train the Human0 flow-matching policyPaper
NS-iHOS / WISH2025-09arXivWeakly supervised in-hand object segmentation from egocentric narrations, distilled for inference without narrationPaper
INSIGHT2025-08AAAI 2026Intention-guided cognitive reasoning for long-term action anticipation using hand-object cues and verb-noun dependenciesPaper
ECHO2025-08arXivRecovers human pose, object motion, and contact dynamics from sparse smart-glasses and wrist-tracker signalsPaper
Ego 6DoF Object Trajectories2025-06CVPR 2025Generates 6DoF object manipulation trajectories from action descriptions in egocentric vision, evaluated with HOT3DPaper
EgoAdapt Efficient Perception2025-06arXivAdaptive multisensory distillation for efficient action, speaker, and behavior anticipation across EPIC-KITCHENS, EasyCom, and AEAPaper
EVA02-AT2025-06arXivEgocentric video-language model family with spatial-temporal rotary positions and symmetric multi-similarity optimizationGitHub
H2R2025-05arXivHuman-to-robot data augmentation that converts egocentric human videos into robot-centric pre-training dataPaper
MEgoHand2025-05arXivMultimodal egocentric hand-object motion generator (VLM + flow-matching) over 3.35M RGB-D frames, 24K interactions, 1.2K objectsPaper
EgoZero2025-05arXivTrains robot manipulation policies from Project Aria smart-glasses human demos with zero robot training dataPaper
Ego4o2025-04CVPR 2025Omni-modal egocentric human motion capture from headset/glasses/phone/watch inputs, sparse IMUs, and motion-language descriptionsPaper
EgoH42025-04arXivDiffusion transformer forecasting both-hand 3D trajectories and poses from egocentric video, including out-of-frame hands (Ego-Exo4D)Paper
Helios 2.02025-03ECCV 2026 Event-Based Multimodal Vision WorkshopSynthetic-only event-camera smart-glasses gesture recognition at 6-8 mW, with more than 80% F1 for its six-channel modelPaper
Fish2Mesh2025-03ICCV 2025Fisheye-aware transformer for 3D human mesh recovery from egocentric vision with weak supervision from third-person camerasPaper
EgoSplat2025-03arXivOpen-vocabulary egocentric scene understanding with language-embedded 3D Gaussian splatting (Aria Digital Twin)Paper
mmEgoHand2025-01arXivHead-mounted millimeter-wave radar and IMU system for egocentric hand pose estimation and gesture recognitionPaper
ObjectRelator2024-11ICCV 2025 HighlightCross-view ego-exo object relation and correspondence model with multimodal fusion and self-supervised alignmentPaper
TouchInsight2024-10UIST 2024Detects touch moment, finger, and location on arbitrary physical surfaces from egocentric hand trackingPaper
EgoZAR2024-09Pattern Recognition LettersZone-aware egocentric action-recognition method improving cross-domain transfer with activity-centric zone priorsProject
3D-Aware Ego Instance Tracking2024-08ACCV 20243D-aware instance segmentation and tracking that lifts 2D masks with scene geometry to survive motion and occlusion (EPIC Fields)Paper
EgoChoir2024-05NeurIPS 2024Infers 3D contact and object affordance from egocentric video, head motion, and 3D object cuesProject
Diff-IP2D2024-05IROS 2025Non-autoregressive diffusion forecasting of 2D hand trajectories and object affordances with camera-egomotion conditioningGitHub
Spatial Cognition / LMK2024-04arXivTracks active objects out of sight by lifting, matching, and preserving 3D object tracks across long EPIC-KITCHENS videosPaper
EffHandEgoNet2024-04FG 2024Egocentric 2D hand pose and action-recognition method for smart-glasses RGB input on H2O and FPHAPaper
EgoLifter2024-03ECCV 2024Open-world 3D segmentation decomposing natural egocentric video into individual 3D objects via 3D Gaussians and SAMGitHub
X-MIC2024-03ECCV 2024Cross-modal instance conditioning for egocentric action generalization across EPIC-KITCHENS, Ego4D, and EGTEAGitHub
EgoPoseFormer2024-03ECCV 2024Transformer baseline for stereo egocentric 3D human pose estimationGitHub
GPT4Ego2024-01arXivFine-grained concept-description prompting for zero-shot egocentric action recognitionPaper
Get a Grip2023-12arXivReconstructs stable hand-object grasps from egocentric videoProject
AV-CONV2023-12CVPR 2024Audio-visual conversational graph prediction from ego/exo conversationProject
Aria-NeRF2023-11arXivMultimodal egocentric view synthesis for Project Aria-style capturePaper
Egocentric Whole-Body MoCap2023-11CVPR 2024FisheyeViT plus diffusion refinement for egocentric whole-body motion capturePaper
Object-Centric LTA2023-10WACV 2024Uses object prompts and visual-language pretrained features for long-term action anticipation on Ego4D and EGTEA Gaze+Paper
Symbolic Active Object Localization2023-10EMNLP 2023Uses symbolic world knowledge to localize active objects from egocentric vision and task instructionsPaper
EgoPCA2023-09ICCV 2023Framework for egocentric hand-object interaction understanding using hand/object-centric cuesPaper
Ego3DPose2023-09SIGGRAPH Asia 2023Captures 3D body cues from binocular egocentric views for pose estimationPaper
Open-Vocabulary Egocentric Actions2023-08NeurIPS 2023Open-vocabulary egocentric action recognition for verb-noun action labelsProject
Helping Hands2023-08ICCV 2023Object-aware egocentric video-recognition model using hand-object cuesPaper
EgoPoser2023-08ECCV 2024Real-time egocentric pose estimation from sparse and intermittent HMD/controller observationsProject
EgoHandTrajPred / USST2023-07ICCV 2023Egocentric 3D hand trajectory forecasting over H2O and EgoPAT3D-style settingsProject
AntGPT2023-07ICLR 2024Uses large language models to improve long-term action anticipation from video contextPaper
AV-EgoGaze Anticipation2023-05ECCV 2024Audio-visual model for forecasting future gaze in egocentric videoPaper
StillFast2023-04CVPRW 2023End-to-end short-term object interaction anticipation for Ego4DProject
EgoViT2023-03arXivPyramid video transformer with dynamic class-token generation for egocentric action recognitionPaper
Next Active Object Anticipation2023-02IEEE Access 2024Anticipates next-active-object locations before contact in egocentric videoPaper
TransFusion2023-01IEEE TCSVT 2024Summarizes past egocentric context in language to improve multimodal object-interaction anticipationPaper
Ego-Only2023-01ICCV 2023Egocentric action detection without exocentric pretraining transferPaper
EgoSTARK2023-01NeurIPS 2023Adapted long-term tracker baseline for EgoTracksPaper
EgoLoc2022-12ICCV 2023Stronger Ego4D visual-query 3D object localization with camera-pose and localization baselinesPaper
CONE2022-11ECCV 2022 WorkshopCoarse-to-fine alignment framework for Ego4D Natural Language QueriesPaper
InternVideo-Ego4D2022-11Ego4D Workshop 2022Pack of champion Ego4D challenge solutions spanning episodic memory, forecasting, hand-object, and audio/social tracksPaper
UmeTrack2022-10SIGGRAPH Asia 2022Unified multi-view end-to-end 3D hand tracking for VR headsets and hand-interaction systemsPaper
EgoHOS model2022-08GitHubContext-aware hand-object segmentation and augmentation pipelineGitHub
I-CVAE LTA2022-07WACV 2023Intention-conditioned variational model for long-term human egocentric action forecastingProject
SOS Handled Objects2022-04ECCV 2022Self-supervised learning over handled-object sets for egocentric action recognitionPaper
HOI Forecast / Interaction Hotspots2022-04CVPR 2022Forecasts future hand trajectories and interaction hotspots on next-active objectsProject
EgoGAN2022-03ECCV 2022Generative future hand-mask forecasting from egocentric videoPaper
Untrimmed Action Anticipation2022-02ICIAP 2022Reframes egocentric anticipation for untrimmed first-person streamsPaper
OWL2022-02CVPRW 2023Audiovisual temporal context model for egocentric action localization on EPIC-KITCHENS and HOMAGEPaper
Hand-Object Interaction Reasoning2022-01AVSS 2022Transformer interaction unit for reasoning over acting hands, the other hand, and interacted objects in egocentric videoPaper
E2(GO)MOTION2021-12CVPR 2022Motion-augmented event-stream representation for egocentric action recognitionPaper
Human Hands as Probes2021-12CVPR 2022Uses hand observations in EPIC-KITCHENS to learn interactive object states, interaction regions, and grasp affordancesPaper
Temporal-Context Ego Action Recognition2021-11BMVC 2021Multimodal transformer that uses surrounding temporal context for egocentric action recognitionPaper
NeuralDiff2021-103DV 2021Neural rendering factorization for segmenting moving 3D objects from dynamic egocentric videoPaper
Head-Motion SSL2021-10ICCV 2021 EPIC WorkshopSelf-supervised representation learning by matching egocentric video clips with AR/VR head-motion IMU signalsPaper
Exo-to-Ego Video Synthesis2021-07ACM MM 2021Cross-view synthesis model that generates egocentric video from exocentric videoPaper
KinPoly / Dynamics-Regulated Kinematic Policy2021-06NeurIPS 2021Object-aware kinematic policy for egocentric pose estimation with scene dynamicsProject
RNA Cross-Domain FPV AV Recognition2021-06arXivRelative Norm Alignment for cross-domain first-person audio-visual action recognitionPaper
Ego-Topo 3D Map Localization2021-05ECCV 2022Joint egocentric activity recognition and localization on a 3D environment mapPaper
Ego-Exo representation transfer2021-04CVPR 2021Distillation from third-person video using ego-specific latent signalsPaper
Egocentric Pose from Human Vision Span2021-04ICCV 2021Egocentric body-pose estimation from a wider glasses-like human vision spanPaper
HPS2021-03CVPR 20213D human pose and self-localization in large scenes from body-mounted sensors and a head-mounted cameraPaper
Indoor Future Person Localization2021-03IROS 2021Future person-location and trajectory prediction from egocentric wearable-camera videoPaper
Contact Representations for Action Forecasting2021-02T-PAMI 2021Forecasts first-person actions through making/breaking hand-object contact representationsPaper
Imagination-Based Ego Action Anticipation2021-01IEEE TIP 2021Anticipates egocentric actions by imagining future visual representationsPaper
RHOI2020-12ICCV 2021Reconstructs hand-object interactions in the wild without direct 3D supervisionProject
Egocentric 4D Human Body Capture2020-113DV 2021Reconstructs second-person 3D human meshes from monocular egocentric video and scene groundingPaper
SelfPose2020-11T-PAMI 20203D egocentric body-pose estimation from downward-looking headset-mounted fisheye camerasPaper
Ego-OMG2020-06arXivObject Manipulation Graph representation for activity modeling and near-future action anticipationPaper
In the Eye of the Beholder2020-05IEEE TPAMIStudies how gaze and wearer attention support action understanding in first-person videoPaper
Unifying Few- and Zero-Shot Ego Action Recognition2020-05CVPR 2020 EPIC WorkshopCompositional few- and zero-shot egocentric action recognition for EPIC-style verb-noun labelsPaper
First-Person Motion-Appearance SSL2020-02ICPR 2020Self-supervised joint motion and appearance encoding for first-person action recognitionPaper
Symbiotic Attention2020-02AAAI 2020 OralUses privileged hand/object information through symbiotic attention for egocentric action recognitionPaper
EGO-TOPO Environment Affordances2020-01CVPR 2020Builds topological maps of interaction opportunities from egocentric videoPaper
Hands in Egocentric Vision Survey2019-12IEEE TPAMIFocused survey of hand localization, interpretation, applications, and major hand-annotated egocentric datasetsPaper
Motor Attention HOI Forecasting2019-11arXivJoint prediction of motor attention and future actions in first-person human-object interactionPaper
Seeing and Hearing Egocentric Actions2019-10ICCV 2019 EPIC WorkshopAudio-visual study and baseline for egocentric human-object action understandingPaper
EPIC-Fusion2019-08ICCV 2019Audio-visual temporal-binding model for egocentric action recognitionPaper
Few-Shot First-Person Action Recognition2019-07T-PAMIDomain-specific priors and meta-learning for few-shot first-person action recognitionPaper
Rolling-Unrolling LSTMs2019-05ICCV 2019 OralMulti-scale LSTM with modality attention for anticipating egocentric actions and interacted objectsPaper
EPIC-KITCHENS action models2019-01GitHubPublic baseline models for EPIC-KITCHENS action recognitionGitHub
Visual Context Event Segmentation2018-08ACM MM 2018Unsupervised segmentation of egocentric photostreams into contextual events, with EDUB-Seg labelsPaper
Task-Dependent Gaze Transition2018-03ECCV 2018Learns task-dependent attention transitions for gaze prediction in egocentric videoPaper
Future Person Localization2017-11CVPR 2018Predicts where people will appear in future frames from a first-person wearable cameraPaper
First-Person Activity Forecasting2016-12ICCV 2017 OralDARKO online inverse-reinforcement-learning system for forecasting actions and active objectsPaper
Unsupervised Important Objects2016-11ICCV 2017Detects important first-person objects without wearer labels using cross-pathway visual/spatial supervisionPaper
Ego-Surfing2016-06IEEE TPAMILocalizes a person in first-person videos using ego-motion signature correlation across wearable-camera viewsPaper
First-Person Pose Recognition2015-06CVPR 2015Uses egocentric workspaces and camera geometry for first-person human pose recognitionPaper
Egocentric FOV Localization2015-01WACV 2015Localizes the camera wearer's first-person field of view in overhead/surveillance videoPaper

Practical Tooling

ToolReleasedVenueUse
Ego-OSCAR2026-08arXivSub-$200 stereo-inertial head-mounted capture design with recording, synchronization, and watchdog software; the paper also claims a 550-hour-per-camera annotated corpus, but no artifact URL is currently linked.
LightMem-Ego2026-07arXivMIT-licensed streaming visual-audio memory for smartphones and Rokid AI glasses, with hierarchical retrieval for object finding, conversation recall, summaries, routines, and evidence-grounded answers.
OpenGlass Visual Assistance2026-07ACL 2026 System DemonstrationsLocal-first ESP32 smart-glasses sensing with nearby-device MLLM inference, speech output, and public evaluation assets.
OpenGlass Event Eyewear2026-06arXivLow-power open AI eyewear platform with event-based vision support for on-device smart-glasses research.
EgoKit2026-05arXivLow-cost synchronized ego/wrist recording workflow across phones, smart glasses, and XR hosts.
GBAT2026-05arXivAnnotating egocentric eye-tracking and video in child-caregiver interaction studies.
MobileEgo Anywhere2026-05arXivCommodity-phone STERA tooling plus a gated 200-hour, 584-session release with 10M frames, pose, IMU, MANO hands, language, and about 1.6 TB of data.
VisionClaw2026-04arXivAlways-on wearable AI-agent system on Ray-Ban Meta smart glasses, coupling live egocentric perception with speech-driven task execution.
HOMIE-toolkit2026-03GitHubLoading and visualizing Xperience-10M HDF5 annotations, calibration, SLAM, hand/body mocap, depth, IMU, and point clouds.
RaycastGrasp2025-10ICIR 2025Egocentric gaze-guided MR headset interface for robotic object retrieval and manipulation.
AiGet2025-01CHI 2025Smart-glasses assistant for gaze/context/profile-driven informal learning during everyday moments.
HUX2024-07arXivAlways-on smart-glasses / XR companion concept with gaze, environment context, verbal context, and memory storage.
HOT3D tooling2024-06project pageLoading HOT3D hand/object/camera pose annotations and models.
Ego-Exo4D CLI and docs2023-11project pageDownloading synchronized ego-exo data and annotations.
EgoObjects API2023-09ICCV 2023Working with category and instance-level egocentric object labels.
EgoBlur2023-08arXivPrivacy-preserving blur pipeline and responsible-innovation analysis for Project Aria egocentric capture.
Project Aria Tools2023-08arXivReading Aria VRS, calibration, MPS, trajectory, gaze, and dataset artifacts.
VISOR API2022-09GitHubLoading dense EPIC-KITCHENS hand/object masks and relations.
HOI4D tooling2022-03project pageLoading RGB-D frames, point clouds, object meshes, and pose/segmentation annotations.
Ego4D CLI and docs2021-10project pageDownloading and working with Ego4D data after license approval.
PAL2021-05CVPR 2021 EPIC WorkshopWearable personalized visual-context detection for privacy-preserving intelligence augmentation.

Adjacent and Related Resources

These resources sit outside the atlas's human-worn egocentric core, but provide useful comparisons and transfer paths: robot-mounted or simulated first-person views, general human video, multi-view robotics, autonomous driving, and long-video reasoning. They are tagged scope: adjacent in `data/resources.yml` so public counts and filters keep the distinction explicit.

ResourceReleasedVenueScale / signalWhy it is adjacent (not egocentric)Status
AlloEgo-VLM2026-08arXivPublic AlloEgo-View triplets and SFT code disambiguate allocentric versus egocentric spatial references, including an Isaac Sim object-search testGeneral images and simulation rather than wearable capture; repository-wide license is absentpartial
WorldRover-10M2026-08arXivSynthetic minute-scale exploration from first-, third-, and 360-degree views with RGB, depth, poses, actions, flow, and long-range tracksArtist-built synthetic worlds; code and data links are not yet livewatch
EATR-Stereo2026-08arXivEmbodiment-conditioned routing of paired head-stereo evidence reports 60% full-task and 100% grasp success on a 33-DoF humanoidRobot head cameras rather than human wearable videowatch
Hydra-02026-08arXivAction-flow world model trained after filtering 2,202 hours of multi-embodiment video, supporting replay evaluation and human-demo transferGeneral robot/human video; code and models are still coming soonwatch
DECOWAM2026-08arXivSeparates camera ego-motion, base, and arm action in a whole-body world-action model and introduces the ARMDOG real-robot datasetMoving robot-camera perception; artifacts are not yet releasedwatch
RoboEdit2026-08arXivTurns human manipulation video into aligned robot video and 3D states; reports RoboEdit-14M with 174K pairs across seven embodimentsHuman video need not be wearable first-person; data and code remain unreleasedwatch
HiPHI2026-08Hugging FaceGated 617.5-hour optical-motion corpus with 200.1M frames, 132 performers, and 245.7 hours of hand-object interactionFull-body and HOI mocap without a central first-person camerarequest
NVIDIA Video-to-Data Robot Dexterity Dataset2026-08Hugging Face1,215 TACO-derived motions retargeted into 121,500 bimanual robot episodes totaling 924.8 GB, with CC-BY-4.0 data and public toolsSynthetic robot executions derived from existing human demonstrationsopen
RoboLab-EgoX2026-08Hugging Face4,000 simulation episodes over 28 tasks and 99 scenes, including success/failure labels, wrist RGB-D, exterior RGB-D, telemetry, and actionsRobot wrist-view simulation benchmark rather than human wearable captureopen
OccPlanner2026-08arXivOccupancy-conditioned pixel-goal diffusion planner with L3ROcc supervision from monocular robot videos; reports 71.55% average success at 5-8 m versus 20.81% for NavDPRobot-view navigation rather than human wearable capturewatch
AdvDex2026-08arXivAligns human and robot dexterous demonstrations through an SE(3) wrist plus 15-joint hand action space, tactile/kinematic OmniShare data, and adversarial embodiment learningOmniShare is not established as a wearable first-person corpuswatch
ContactGuard2026-08arXivPredicts action-conditioned visual latents and aborts likely manipulation failures before contact from wrist and multi-view robot camerasRobot execution monitoring rather than human wearable capturewatch
HumanoidVLN2026-08arXivIsaac Sim benchmark with four humanoids, 933 collision-aware routes, four instructions per route, and a 20-episode Unitree G1 sim-to-real pilotSimulated humanoid first-person observations rather than human wearable capturewatch
Proxemic Risk from Robot Ego Images2026-08ECCV 2026 EMR WorkshopFour-level danger classification study of InternVL, Qwen-VL, and SmolVLM with prompting, QLoRA, and person-localization analysisRobot viewpoint and safety labels rather than human wearable capturewatch
Fast-Slow ReAct ObjectNav2026-08arXivBounded VLM deliberation over coordinate-anchored semantic memory and pose-tagged first-person keyframes; reports 68.75% HM3D and 47.29% MP3D successNavigation-robot observations rather than human wearable capturewatch
VideoNIG2026-08arXiv60K simulated ego-view tours and 37K multimodal prompts for navigation-instruction generation, spatial diagnosis, and downstream executionSimulator-generated agent views rather than human wearable capturewatch
WNM-3D2026-08arXivConditions joint future-view and action generation on geometry tokens reconstructed from monocular navigation historyNavigation-agent camera history rather than human wearable capturewatch
AtlasVLA2026-08arXivWrist-camera VLA with persistent voxel-hashed 4D world memory and task-progress-aware ego working memoryRobot wrist-camera manipulation rather than human wearable capturewatch
CrossTracer2026-08arXivCross-embodiment image-plane trace prediction and residual adaptation, evaluated on NaviTrace and wheeled/legged robotsRobot-mounted egocentric sensing rather than human wearable capturewatch
Omega-02026-08arXivWhole-body latent world-action model supporting robot ego RGB; reports Omega-HOME with 40+ hours of synchronized household humanoid observations and actionsRobot-mounted egocentric sensing rather than human wearable capturewatch
LAWM-3D2026-08arXivLearns 3D-aware latent actions from multiview human video for generalizable robot world modelsSource human video is not confirmed as wearable first-person capturewatch
HOST2026-07arXivOne-video skill acquisition in 29 seconds on average with 62% reported success; training code and a 15.9 GB policy checkpoint are publicHuman video is central but not required to be wearable first-personpartial
Data Pyramid for Embodied Manipulation2026-07arXivOpen survey and living catalog organizing real-robot, UMI, human ego/ego-exo, simulation, and general video data into five levelsBroad embodied-manipulation taxonomy in which egocentric data is one major layeropen
ContactFlow2026-07arXivEmbodiment-agnostic world-model conditioning through trajectories of 3D actor-object contact points, trained with human and robot interaction videoGeneral human/robot video; wearable capture is not requiredwatch
Robot-Factored World Models2026-07arXivFactors actions into nominal robot trajectories and rendered geometry across DROID, egocentric RoboCasa-GR1, unseen robots, and retargeted human demosRobot-camera world model rather than human wearable capturewatch
AWM Navigation Pretraining Corpus2026-07Hugging FacePublic 198.1 GB corpus with 23,076 indoor/outdoor robot-view clips and paired Gemini-generated spatial/action labels; source-derived reuse terms are incompleteRobot navigation observations rather than human wearable capturepartial
Think at 5 Hz, Act at 20 Hz2026-07arXivAsynchronous driving VLA with a frozen 7B reasoner and fast action expert, raising CARLA route completion from 37.0 to 94.0 at fresh 20 Hz controlAutonomous-driving camera input rather than human wearable capturewatch
IMBench2026-07RSS 2026 SemRob WorkshopOpen benchmark with 35 tasks, seven categories, and 14K trajectories integrating physical reasoning with executable manipulationRobot-manipulation benchmark rather than human wearable captureopen
AC-VLA2026-07arXivCompositional VLA training with instruction decomposition, trajectory alignment, and wrist-view masking; reports about 28% absolute LIBERO-OOD improvementRobot VLA method rather than a human first-person resourcewatch
SkillNav2026-07arXivTraining-free object-goal navigation skills written into a curiosity map, reaching 43.2 SPL and 75.9% success on HM3D v0.2Robot navigation observations rather than human wearable capturewatch
Orbis 22026-07arXivHierarchical driving world model that separates long-horizon scene structure from detailed generation; code and checkpoints remain coming soonAutonomous-driving world model rather than human wearable capturewatch
Data and Learning Where it Matters2026-07arXivTargeted collection and offline RL for contact-critical segments averages 96% success across four real tasks from 2-2.5 hours of dataRobot contact-rich manipulation rather than human wearable capturewatch
MotionForesight2026-07arXivForecasts future object-centered 3D scene flow by repurposing frozen video and tracking priors trained on 40K monocular human-object videosGeneral monocular interaction video without a wearable-capture requirementwatch
VTM-Nav2026-07arXivTraining-free object-goal navigation with persistent visual-topological room/object memory, coarse-to-fine retrieval, and a conservative execution guardRobot navigation memory rather than human wearable capturewatch
SafeRelBench2026-07arXiv507 household-agent evaluations testing support, containment, and proximity constraints before risk-prone actionsRobot-agent spatial-safety evaluation rather than human wearable capturewatch
SoftNav2026-07IROS 2026Continuous 3D soft tokens ground a frozen VLM in detected objects and frontiers using about 1,200 samples and 17M trainable parametersRobot object-goal navigation rather than human wearable capturewatch
Representation-Aligned Tactile Grounding2026-07arXivLatent tactile predictor aligns intermediate VLA features with future contact consequences without reconstructing noisy raw tactile signalsRobot tactile-VLA learning rather than human wearable capturewatch
Action QFormer2026-07arXivInstruction-conditioned query interface reorganizes inherited multimodal features before action generation, with large reported sim-to-real gainsRobot navigation VLA architecture rather than human wearable capturewatch
Reflex2026-07ICML 2026Streaming flow-matching VLA runtime reports 2.58x speedup, stable 50 Hz control, and up to 54% lower reaction latencyRobot VLA inference framework rather than human wearable capturewatch
FoMoVLA2026-07arXivFuture-feature foresight and sparse 2D trajectories jointly supervise target states and motion paths for continuous policiesRobot VLA foresight method rather than human wearable capturewatch
LifelongVLA2026-07arXivContinual VLA learning separates reusable shared knowledge from task-specific experts while controlling forgettingRobot continual-learning method rather than human wearable capturewatch
Steering Robustness into World Action Models2026-07arXivRobustness-oriented WAM training uses uncertainty-aware perturbations and corrective trajectories to improve recovery under distribution shiftRobot world-action-model training rather than human wearable capturewatch
AeroAct2026-07arXivVision-language-action policy couples aerial perception, language-conditioned planning, and low-level flight controlAerial-robot VLA rather than human wearable capturewatch
DriftWorld2026-07arXivDriving world model explicitly represents scene dynamics and controllable ego motion for long-horizon generationAutonomous-driving world modeling rather than human wearable capturewatch
BadWAM2026-07arXivBenchmark and attack framework studies backdoors in world-action models across visual triggers and action outcomesRobot WAM security rather than human wearable capturewatch
RoboTTT2026-07arXivOne-shot human-video imitation combines online test-time training with world-model supervision, reporting five-minute learning of a 10-stage assembly taskHuman-video robot adaptation without a released wearable-video corpuswatch
PhysClaw-02026-07arXivOpen collect-verify-reset system that stores and reuses language corrections, reducing human working time to 16% and raising single-attempt success from 12.5% to 47.5%Robot autonomous-data-collection toolkit rather than human wearable captureopen
Industrial Dexterity Benchmark2026-07arXivIndustrial hardware boards plus DAG-ROS and AG-iDP3 for cable, harness, and gearbox tasks; best cable setup reaches 78% versus 36% for single-camera RGBIndustrial robot benchmark rather than human wearable capturewatch
GigaWorld-Policy-0.52026-07arXivOpen action-centered WAM with mixed world/action training, MoT experts, AutoResearch, and 85 ms RTX 4090 inference without future-video decodingRobot world-action model rather than human wearable videoopen
VSI-Super-Wild2026-07ECCV 20266,980 human-verified spatial QA pairs over 442 continuous panoramic videos totaling 284.52 hours; public 2.67 TB data lacks explicit reuse termsPanoramic YouTube-derived video rather than wearer capturepartial
REAL / REAL-Bench2026-07ECCV 2026Open mobile-manipulation framework and 241-task benchmark reporting 56.9% interactive and 78.3% physical-robot success, with code and 3.92 GB of dataRobot-only open-world manipulation rather than human wearable captureopen
UESF-Bench / SeekFollow-VLA2026-07arXivUnified language-guided human seeking and following with semantic exploration, delayed identity grounding, phase switching, and recoveryRobot-agent human search/following rather than human wearable videowatch
JOP-VLN2026-07IROS 2026Joint imitation, DAgger, and reinforcement learning reaches 69.9% success on R2R and 68.0% on RxRNavigation-agent observations rather than human wearable capturewatch
Reverse to Advance2026-07arXivTrains difficult manipulation from easier reversed rollouts using automated collection, temporal inversion, kinematic filtering, critic filtering, and iterative refinementRobot manipulation data collection rather than human first-person datawatch
Anchor-Align2026-07arXivRepresentation anchoring plus language-action alignment improves physical xArm success from 28% to 54% and 37% to 60%; linked code currently returns 404Robot-camera VLA finetuning rather than human wearable capturewatch
JITOMA / JITOMA-Bench2026-07arXivJust-in-time scene-graph growth activates task-relevant memory anchors to curb long-horizon perceptual saturation, graph size, and caption latencyRobot scene-graph memory rather than human wearable videowatch
WANDA / Worlds in One Demo2026-07arXivPublic 16,922-episode, roughly 248-hour synthetic mobile-manipulation set expands one real demonstration per task across generated 3D worldsSynthetic robot trajectories rather than human wearable captureopen
GPUSimBench2026-07IROS 2026Tests Isaac Lab and Genesis for real-to-sim physical consistency, parallel throughput and memory, plus run-to-run and inter-environment nondeterminismRobot-simulator infrastructure benchmark without human capturewatch
HRIBench (Interaction-Centric VLA)2026-07arXiv13 role-conditioned tasks and 650+ episodes assess intent, synchronization, responsiveness, protocol, and safety across instructor, collaborator, and intruder rolesRobot-view collaboration benchmark rather than human wearable capturewatch
DenseReward2026-07arXivDense frame-level vision-language rewards trained with automatically synthesized physical failures for manipulation, MPC, and RL; the claimed release page currently returns 404Robot-camera reward learning rather than human wearable capturewatch
FlowWAM2026-07arXivOpen optical-flow world-action model reporting 92.94%/92.14% RoboTwin Clean/Random success and a 63.71 WorldArena EWMScore, with code, checkpoints, and flow dataRobot-only manipulation and world modeling rather than human wearable captureopen
ChunkFlow2026-07arXivContinuity-consistent chunked-policy learning with editable overlap zones, seam and derivative losses, robust history training, and AWAC adaptationRobot-policy stability method without a human first-person resourcewatch
ExToken2026-07arXivReinforcement fine-tuning that conditions VLA exploration on demonstration-derived behavioral tokens and learns a deployment-time token selectorRobot VLA exploration method rather than egocentric human datawatch
Hy-Embodied-VLM-1.02026-07arXivOpen approximately 30B MoE embodied VLM activating about 3B parameters per token; evaluated on 38 benchmarks with first-place results on 19 among compared similarly sized modelsGeneral physical-world agent model rather than a human wearable resourceopen
Jetson-PI2026-07arXivOpen foresight-aligned asynchronous VLA stack reporting 8.66x higher control frequency than naive PyTorch, 5.41x over vla.cpp, and +14.8% average success over VLASHOnboard robot-control runtime rather than human first-person captureopen
TrustVLA2026-07arXivRetraining-free VLA backdoor defense using evidence-evolution monitoring, counterfactual trigger localization, and localized inpaintingRobot VLA security method rather than an egocentric data resourcewatch
Self-in-Space / SIS-Bench2026-07ACM MM 2026Open 4,856-question benchmark over 1,646 UAV videos and 13 tasks, plus 54,298 training examples, code, and a motion-aware adapterAgent-centered UAV onboard video rather than human wearable captureopen
VistaVLA2026-07arXiv3D-Gaussian-grounded VLA reporting 99% context-token reduction, +22.8% success across seven real tasks, and +30.0% OOD over VLA-AdapterMulti-view robot manipulation rather than human wearable videowatch
Efficient VLA Inference via Temporal Redundancy Reduction2026-07arXivDynamic visual-token updates plus two-step diffusion action sampling report more than 2x speedup and up to 98% success across LIBERO, RoboTwin, and real robotsRobot VLA inference optimization rather than a first-person resourcewatch
REGRIND2026-07arXivSingle-demonstration retargeting plus residual RL for scissors and screwdriver use on LEAP and WUJI hands; linked code repository is currently unavailableHuman hand-object references and robot control, but no wearable-video releasewatch
E-VQA / ST-Evidence2026-07ECCV 2026Open benchmark for answers with temporal segments and tracked object masklets, plus 160K instruction examples and released 3B/7B modelsGeneral video QA and evidence grounding rather than egocentric videoopen
Xiaomi-Robotics-U02026-07arXiv38B embodied-synthesis model for multi-view scenes, embodiment transfer, editing, and video; the claimed release URL currently returns 404Robot and embodied-scene generation rather than human wearable capturewatch
World Action Models to Embodied Brains Roadmap2026-07arXivRoadmap spanning WAM representations, standardization, physical harnesses, shared contracts, and closed-loop post-trainingBroad physical-intelligence position paper rather than an egocentric resourcewatch
DA-Nav / ReDA2026-07arXivDirection-aware city-scale VLN and recovery data, reporting 56.16% unseen-city success and kilometer-scale robot trialsRobot-agent egocentric navigation, not human wearable capturewatch
Robot-Centric Pointmaps2026-07arXivRobot-frame per-pixel XYZ improves pi0.5 by 7.6 points and SmolVLA by 4.2 on RoboCasa, with larger gains under unseen real-camera viewpointsRobot-camera VLA representation rather than human first-person datawatch
TeleDexter2026-07arXivHand-object co-tracking controller reporting 75.2% average success across seven tasks and two dexterous hands, with demonstrations used for behavior cloningRobot teleoperation from human references, not a released wearable-video resourcewatch
WALA2026-07arXivExecutable latent actions from labeled demonstrations and action-free videos using DINOv3 feature deltas, depth, and latent world modelingGeneral robot-policy learning; code is still marked coming soon and no egocentric corpus is releasedwatch
SLVMBench2026-07arXiv2,261 questions over 1,220 videos test learning a tutorial embedded in a 2-3-hour stream and applying it to an ongoing taskGeneral long-video skill memory rather than first-person-centered datawatch
Lumo-22026-07arXivLatent WAM with staged action-dynamics-vision-language alignment for scalable, OOD, long-horizon, and dexterous robot learningRobot-centered WAM/VLA method without a wearable-video releasewatch
AdvNav2026-07ACM MM 2026Black-box adversarial attack on first-person VLN observations with 49.70%-87.30% reported attack success against HAMT and MapGPTNavigation-agent views rather than human wearable videowatch
yyyyywv/egocentric2026-07Hugging Face32.8 GB, 1,019 robot episodes, and 994,459 frames across seven task groups from eye/head and wrist cameras; downloadable but the card omits provenance and capture documentationRobot-mounted cameras and action/state streams rather than human wearable capturepartial
TactiDex2026-07arXivHuman-demonstration benchmark aligning whole-hand tactile pressure, hand kinematics, object 6D states, language, and task phases, plus tactile-guided single/bimanual robot transferTactile-glove and motion-capture demonstrations, not human wearable videowatch
DemoBridge2026-07RSS 2026 RoboData WorkshopApache-2.0 toolkit that converts single-view stereo hand demonstrations into collision-aware, physics-validated robot trajectoriesHuman-hand retargeting without explicitly egocentric captureopen
PhysV2A2026-07arXivConverts video-derived 6D object motion into feasible robot trajectories using grasp-conditioned reachability checks and semantic-mask-constrained refinementGeneric video-to-robot transfer rather than human wearable capturewatch
ACE-Brain-0.52026-07arXivPublic Qwen3-VL-derived 8B-class checkpoint unifying spatial perception, planning, robot action generation, progress monitoring, and self-improvement across 15 benchmarks; no model license is declaredRobot-centric embodied foundation model with egocentric spatial reasoning, not human wearable capturepartial
DexVerse2026-07arXiv100-task dexterous-manipulation benchmark across three arms and six hands, with VR teleoperation and 3,180 reported demonstrations; core BSD-3-Clause suite is live while full data/tooling remain pendingRobot simulation and teleoperation benchmark, not a human wearable-video corpuspartial
FabriVLA2026-07arXivCompact 1B VLA with InternVL3.5 and a flow-matching action head, reporting 90.0% tier-average success on Meta-World MT50Simulation-focused robot VLA, not human egocentric capturewatch
Harness VLA2026-07arXivMemory-guided agentic harness composing a frozen VLA with retryable contact primitives and fixed analytic primitives across LIBERO-Pro, RoboCasa365, and RoboTwin C2RRobot VLA planning harness, not a first-person resourcewatch
AnyDexRT2026-07arXivCalibration-free hand retargeting with self-supervised fingertip correspondence, few-shot human guidance, and contact-aware pinch refinementHuman-hand teleoperation method, not egocentric video datawatch
TFP2026-07RSS 2026 SemRob WorkshopEvent-sensitive memory-action fusion for a 3.3B VLA, evaluated on LIBERO, LIBERO-plus, and MIKASA ShellGameTouchRobot VLA memory method without wearable capturewatch
LingBot-VA 2.02026-07arXivNative video-action model with a semantic tokenizer, causal pretraining, sparse MoE backbone, and asynchronous closed-loop inferenceRobot video-action foundation model, not human first-person datawatch
LEEVLA2026-07arXivTask-aware VLA using drift-guided visual prioritization and structured latent feature-flow generation; the linked repository does not yet match the paperRobot VLA architecture without human wearable capturewatch
Video-Action Generalization Gap / Temporal Ratio2026-07arXivDiagnostic Temporal Ratio and adaptive inference guidance for compositional generalization in video-action models on LIBERO and real robot tasksRobot video-action analysis, not a human first-person resourcewatch
GIRAF2026-07CVPR 2026 HuMoGen WorkshopText-conditioned full-body interaction generation with articulated objects, object-centric contact representations, and contact augmentationGenerated third-person human-object motion, not wearable capturewatch
LingBot-World 2.02026-07arXivInteractive world simulator with unbounded interaction horizon, 720p/60 FPS real-time distillation, richer action/event controls, and 14B/1.3B modelsGeneral interactive world simulator, not human first-person capturewatch
RoboDojo2026-07arXivUnified sim-and-real robot manipulation benchmark with 42 simulation tasks, 18 real tasks, remote evaluation, 30 integrated policies, and a leaderboardRobot-policy benchmark, not human wearable capturewatch
LaMem-VLA2026-07arXivLatent-memory-native VLA framework that reconstructs short- and long-term history into latent memory tokens for long-horizon manipulationRobot VLA memory architecture, not a human first-person resourcewatch
TouchWorld2026-07arXivPredictive-and-reactive tactile foundation model separating vision-language planning, tactile world prediction, action generation, and residual refinementRobot tactile manipulation method, not wearable capturewatch
WAM-TTT2026-07arXivTest-time training framework steering frozen WAMs from unlabeled human videos through adaptive memory and paired human-robot meta-trainingHuman-video-to-robot WAM method, not a human wearable datasetwatch
EAGOR2026-07arXivTraining-free spherical-belief framework for 360-degree directional reasoning on HOS, OSR-Bench, and legged-robot navigationEmbodied 360-camera robot/agent reasoning, not human wearable capturewatch
VLA Models Review2026-07arXivSurvey of 183 VLA contributions from 2017-2026 across architectures, training recipes, bimanual coordination, UAVs, memory, and world modelsRobot/UAV VLA survey, not an egocentric resourcewatch
NativeMEM2026-07arXivVLA memory compression that turns each historical camera frame into a single memory token for low-latency long-horizon manipulationRobot VLA memory method, not human wearable capturewatch
Lift3D-VLA2026-07arXiv3D geometry- and dynamics-aware VLA with point-cloud reasoning, geometry-centric masked autoencoding, and 22 sim / 8 real tasksRobot-only 3D VLA method, not human first-person datawatch
Pelican-VLA 0.52026-07arXivUnified VLA report combining vision-language understanding, future-frame generation, action prediction, and reasoning-slot attentionRobot VLA architecture, not human wearable capturewatch
ActionCache2026-07arXivTraining-free cache/refinement method accelerating flow-matching VLA action heads by reusing intermediate actions with compact multimodal keysRobot VLA inference method, not a first-person resourcewatch
Smooth Operator / SBR2026-07arXivSampling-based hand retargeter for low-jitter real-time teleoperation, evaluated in simulation and an 18-participant manipulation studyTeleoperation/retargeting method, not a first-person datasetwatch
LongVQUBench2026-07ECCV 20261,200+ long videos and 1,500 questions for video-quality reasoning, spanning movies, documentaries, surveillance, egocentric recordings, and animationBroad long-video quality benchmark with egocentric recordings as one slice, not centered on wearable capturewatch
TAP VLA2026-07ICML 2026Task-agnostic VLA pretraining learns motor priors from unlabeled robot interaction data before language groundingRobot/VLA pretraining, not human wearable capturewatch
The Moving Eye2026-07IROS 2026Hybrid dynamic data collection with one robot arm acting as a moving environmental camera while the other manipulatesRobot camera/viewpoint generalization, not human first-person capturewatch
Bridge-WA2026-07arXivWorld-action framework distilling future-change tokens, change maps, and motion-flow maps into a VLA action transformerRobot world-action model, not wearable capturewatch
VLA-Corrector2026-07arXivLightweight detect-and-correct inference for action-chunked VLA policies using visual feature evolutionRobot VLA correction method, not human wearable capturewatch
Embodied.cpp2026-07arXivPortable C++ runtime for VLA/WAM inference on heterogeneous robots with adapters, head plugins, and robot deployment layersEmbodied robot runtime, not a human first-person resourcewatch
EVA-Client2026-07arXivUnified real-robot client for policy deployment, data collection, evaluation, rollout logging, and comparison viewsRobot policy client and evaluator, not wearable capturewatch
InternVLA-A1.52026-07arXivRobot VLA with native VLM backbone, continuous action expert, subtask/VQA supervision, and latent foresight tokensRobot-only VLA architecture, not human first-person datawatch
SIEVE2026-07arXivStructure-aware selection of VLA demonstrations by reusable primitives, composition patterns, and medoid trajectoriesRobot demonstration-selection method, not egocentric capturewatch
SVA / Look Before You Leap2026-07arXivDistills tree-search rollouts into an action evaluator for frozen VLA policies at test timeRobot VLA action-evaluation method, not wearable capturewatch
WorldOdysseyBench2026-06arXiv600+ first/third-person interactive world-model test cases with 10-60s WASD interaction and action, vision, physics, and memory metricsGeneral interactive world-model benchmark, not human wearable capturewatch
Qwen-RobotNav2026-06arXivQwen-branded navigation model trained on 15.6M samples with task modes and controllable visual-history observation parametersRobot/agent navigation model, not human wearable capturewatch
Qwen-RobotWorld2026-06arXivLanguage-conditioned embodied video world model with an 8.6M video-text EWK corpus over 20+ embodiments and 500+ action categoriesRobot, navigation, driving, and human-to-robot world modeling rather than human first-person capturewatch
FutureNav2026-06arXivUnified world-action modeling for vision-and-language navigation with action, dynamics, and future spatial-state objectivesNavigation-agent WAM, not human wearable capturewatch
ZR-02026-06arXiv2.6B VLA trained with dense embodied chain-of-thought supervision over ProcCorpus-60M for cross-embodiment manipulationRobot VLA pretraining, not a human first-person datasetwatch
VLK2026-06arXiv48K synthetic humanoid loco-manipulation trajectories with rendered egocentric observations, instructions, and whole-body kinematicsSynthetic humanoid supervision, not real human wearable capturewatch
HUMEMBR2026-06IROS 2026Human-centered long-term memory for embodied robot QA and routine-conditioned navigationRobot memory/navigation system, not an egocentric capture resourcewatch
RhinoVLA2026-06arXivEdge-deployable VLA using a token-efficient Qwen3-VL backbone, Action Expert, View Registry, and hardware-aware executionRobot-only VLA deployment system, not wearable capturewatch
OpenEAI-Platform2026-06arXivOpenEAI-Arm plus OpenEAI-VLA built on Qwen3-VL-4B and a Diffusion Transformer action headRobot-arm hardware/software platform, not human first-person datawatch
SAGE-Nav2026-06arXivLLM-planned object-goal navigation with dynamic scene graphs and egocentric robot observationsRobot object-navigation method, not human wearable capturewatch
USS2026-06arXivUnified spatial-semantic prompting for embodied visual tracking with text, point, box, and mask promptsRobot/agent tracking method, not human first-person capturewatch
RoboAtlas2026-06arXivContextual Active SLAM with OpenRoboVox semantic mapping and egocentric VLM reasoningRobot active-SLAM/navigation system, not human wearable capturewatch
iCrowdNav2026-06arXivIntention-aware crowd navigation from egocentric robot observations and human-pose interaction encodingRobot crowd-navigation method, not human wearable capturewatch
Cloak2026-06arXivMasks the robot end-effector from wrist-camera VLA observations for zero-shot cross-embodiment transferWrist-camera robot VLA method, not human first-person capturewatch
OpenHLM2026-06arXivWhole-body humanoid loco-manipulation recipe with teleoperation, VLA design, and HuMI co-trainingRobot-centered humanoid VLA recipe, not human wearable capturewatch
UniviewVLA2026-06arXivMultiview VLA with world modeling for occlusion-aware robot manipulation from agent and wrist camerasRobot VLA/world-modeling method, not human first-person capturewatch
ContactWorld2026-06arXivVision-tactile world-model benchmark over 12 contact-rich manipulation tasks with wrist/front views, point clouds, and tactile force fieldsRobot contact-rich manipulation benchmark, not human wearable capturewatch
3DG-VLN / UAV-VLN-FOV2026-06arXiv2,717 target-visible UAV trajectories with high-resolution egocentric observations and 3D waypoint labelsAerial-agent navigation benchmark, not human wearable capturewatch
NavWAM2026-06arXivNavigation world-action model predicting future egocentric visual observations and action sequences for goal-conditioned visual navigationRobot visual-navigation world model, not human wearable capturewatch
IntentNav2026-06arXivSpatial-visual object-navigation policy learned from human demonstrations with frontier probing and spatial memoryObject-navigation work from human demonstrations, not a human wearable resourcewatch
AlloSpatial2026-06arXivAgentic spatial-reasoning harness that converts local egocentric observations into allocentric spatial representationsGeneral spatial-reasoning framework, not a wearable capture datasetwatch
World-Language-Action Model2026-06arXivUnified model predicting subtasks, subgoal images, and robot actions while learning from egocentric videos and robot trajectoriesBroader embodied model using egocentric video among robot data, not centered on human wearable capturewatch
ManiSplat2026-06arXivManipulation trajectory synthesis from monocular video using decoupled 3D Gaussian SplattingMonocular manipulation synthesis; first-person human capture not confirmedwatch
DexFuture2026-06arXivHierarchical future-state visuomotor targeting for bimanual dexterous tool useRobot dexterity method, not human first-person datawatch
GRAIL2026-06arXivDigital-twin pipeline generating humanoid loco-manipulation demonstrations from 3D assets and video priorsHumanoid demonstration generation, not a human wearable datasetwatch
DataLadder2026-06arXivSimulation-enabled interconversion toolchain for robot trials, demonstrations, synthetic data, and evaluation assetsEmbodied-data tooling rather than egocentric capturewatch
EgoInfinity2026-06arXivPublic 46 GB preview with 104 Action100M clips and retargeted trajectories for four robot embodiments; the full web-scale engine is not releasedHuman-video-to-robot data engine, not centered on wearable first-person capturepartial
LUCID2026-06arXivEmbodiment-agnostic intent model learned from unstructured human videos, paired with simulation-trained dexterous robot controlUnstructured human-video-to-robot transfer; source videos are not confirmed as egocentric wearable capturewatch
Dexterous Point Policy2026-06arXivTransfers raw human videos into dexterous hand policies through task-relevant 3D keypoints for hands and objectsHuman demonstration video transfer, not confirmed first-person capturewatch
TopoRetarget2026-06arXivInteraction-preserving retargeting from human hand-object demonstrations to dexterous robot-hand referencesHuman hand-object retargeting method, not confirmed wearable first-person capturewatch
WARP Retarget2026-06arXivWhole-body-aware retargeting from offline human demonstrations into mobile-manipulator trajectoriesOffline human-demonstration retargeting, not confirmed wearable first-person capturewatch
V2P-Manip2026-06arXivLearns dexterous manipulation from monocular human videos through 3D assets, trajectory estimation, physical refinement, and policy learningMonocular human-video-to-robot work, not confirmed wearable capturewatch
SeeTraceAct2026-06arXivDemo-conditioned VLA with visibility-aware future end-effector traces and RoboCasa-DC cross-embodiment demonstration episodesRobot VLA conditioned on demonstrations, not a human wearable datasetwatch
What Matters When Cotraining Robot Manipulation Policies on Everyday Human Videos?2026-06arXiv532 everyday human videos with 28 hours of high-quality hand labels for studying robot-policy cotraining under limited robot dataEveryday human-video transfer study, not scoped as egocentric capturewatch
LARA2026-06arXivLatent Action Representation Alignment for jointly optimizing latent action models and VLAs from unlabeled human videosHuman-video-to-VLA method, not itself an egocentric datasetwatch
Video2Sim2Real2026-06arXivReconstructs a simulator-ready digital twin and robot/object motion priors from a single human manipulation videoSingle-human-video robot skill acquisition, not confirmed first-person wearable datawatch
HOWTransfer2026-06arXivLocalizes hand-object contact in human video demonstrations and retargets wrist trajectories into robot-executable motionsHand-centric human-video-to-robot transfer, not confirmed egocentric wearable capturewatch
Retrieve, Don't Retrain2026-06arXivRetrieval-augmented VLA adaptation using pool-side demonstrations such as human-hand videos instead of per-task retrainingRobot VLA adaptation with human-hand video pools, not centered on first-person datawatch
SERF2026-06arXivSpatiotemporal environment and robot feature map updated from egocentric observations and proprioception for long-horizon mobile manipulationRobot-centric policy work using egocentric observations, not human wearable capturewatch
ImageWAM2026-06arXivWorld-action model that uses image editing instead of dense video generation for robot action predictionRobot WAM comparison point, not a human first-person resourcewatch
CHORD2026-06arXivNVIDIA contact-wrench-guided dexterous RL with 4,739 simulation tasks, 82.12% success over 1,831 evaluated tasks, whole-body transfer, and real-hardware resultsMotion-capture and third-person human demonstrations rather than a wearable-video release; code remains forthcomingwatch
JAXenstein2026-05arXivJAX Wolfenstein 3D renderer for fast reinforcement-learning experiments in visual first-person tasksVirtual first-person RL environment, not real egocentric capturewatch
ACWM-Phys2026-05arXivControllable action-conditioned world-model benchmark for rigid, kinematic, deformable, and particle dynamicsSimulator benchmark for physical interaction, not human first-person capturewatch
SpatialBench2026-05arXivBroad spatial foundation-model benchmark across viewpoints, scene domains, input densities, and hardware constraintsUseful context for egocentric spatial reasoning, but not centered on human wearable capturewatch
LEXI-SG2026-05arXivMonocular RGB open-vocabulary 3D scene-graph mapping validated on self-collected egocentric office sequencesRobot/agent mapping work, not centered on human wearable capturewatch
EgoTraj-Bench2026-04ICRA 202636,947 public sequences from 17 robot FPV sessions for converting noisy first-person trajectories into clean bird's-eye trajectoriesRobot-mounted first-person views rather than human wearable captureopen
EgoDyn-Bench2026-04ECCV 2026Public 1,000-clip physics-grounded driving benchmark with about 14K oracle-labelled QA pairs, dynamics arrays, and a 49-model reference leaderboardAutonomous-driving ego-motion video rather than human or animal wearable captureopen
DenseStep2M2026-04arXivAbout 100K instructional videos and 2M dense procedural steps from automated long-video annotationGeneral instructional-video corpus with egocentric transfer tests, not first-person centeredwatch
StarVLA2026-04arXivModular VLA codebase supporting swappable VLM backbones such as Qwen-VL and world-model backbones such as CosmosVLA tooling for comparisons, not human first-person capturewatch
Point of View Robot Sociability Study2026-03arXivImmersive-VR study comparing allocentric, egocentric-proximal, and egocentric-distal ratings of identical robot trajectories and a head-nod signalSimulated pedestrian viewpoints rather than recorded wearable videowatch
RoboMirror2025-12arXivRetargeting-free humanoid locomotion from raw egocentric or third-person video using VLM-derived motion intent and a diffusion policySupports egocentric input but targets humanoid locomotionwatch
VLA-Arena2025-12ICML 2026Open-source VLA benchmark with 170 structured tasks and decoupled task, language, and visual perturbation axesRobot VLA benchmark, not a human wearable/egocentric datasetwatch
VLSA / AEGIS2025-12IROS 2026Plug-and-play safety constraint layer for VLA policies plus SafeLIBERO safety-critical manipulation benchmarkRobot VLA safety architecture, not human wearable capturewatch
UniWM2025-10ECCV 2026Memory-augmented world model for navigation foresight and planning, with public code and a Hugging Face datasetRobot navigation cameras rather than human wearable capturepartial
MM-Nav2025-10arXivMulti-view VLA navigation model with 360-degree observations and synthetic expert data for reaching, squeezing, and avoidingRobot/synthetic visual navigation rather than human wearable capturewatch
Seeing Across Views / MV-RoboBench2025-10ICLR 20261.7K curated QA items over eight subtasks for multi-view spatial reasoning of VLMs in robotic manipulation (ICLR 2026)Multi-camera robot scenes, not wearable captureopen
HUI3602025-09FG 202699 mobile-robot 360-degree HRI recordings, a 1M-annotation HUI360 open set, 6M SSUP-HRI annotations, and public anticipation baselinesCaptured from a mobile robot, not a human camera wearerpartial
HomeSafeBench2025-09arXivOpen VirtualHome benchmark with 1,000 test tasks, 3,400 training tasks, five hazard classes, and 3,158 CueBack trajectoriesRendered embodied-agent observations rather than wearable captureopen
RoboPearls2025-06arXivEditable 3D Gaussian video simulation for robot manipulation, evaluated on RLBench, COLOSSEUM, Ego4D, Open X-Embodiment, and real robotsRobot simulation/editing tool, not a human egocentric datasetwatch
NORA2025-04arXiv3B generalist VLA using Qwen2.5-VL-3B and 970K real-world robot demonstrationsRobot-only demonstrations and observations, not human wearable capturewatch
SEED4D2024-12WACV 2025Synthetic ego-exo dynamic 4D generator and autonomous-driving dataset (16.8M images, vehicle cameras, LiDAR; WACV 2025)Vehicle-egocentric driving data, not human/wearableopen
Open X-Embodiment / RT-X2023-10ICRA 20241M+ real robot trajectories, 22 robot embodiments, 60 pooled robot datasets, standardized RLDS; unofficial Hugging Face mirrorRobot-mounted and wrist cameras, not human first-person; pairs with egocentric human-video VLA pretrainingopen

Surveys, Papers, and Context

Foundational papers and recent surveys for background and citation.

ResourceWhy it matters
Vision-Language Models for Egocentric VideoCritical survey connecting hand-object interaction, temporal and graph reasoning, efficient long-video processing, wearable assistance, and embodied AI.
Ego4D: Around the World in 3,000 Hours of Egocentric VideoDefines the modern large-scale egocentric benchmark suite.
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesDefines the major synchronized ego-exo activity resource.
Scaling Egocentric Vision: The EPIC-KITCHENS Dataset and EPIC-KITCHENS-100Core kitchen/action-recognition benchmark lineage.
Egocentric Video-Language PretrainingIntroduces EgoClip and EgoNCE, a central egocentric VLP recipe.
EgoSchemaA widely used diagnostic benchmark for very long-form video-language understanding.
HoloAssistImportant for human-AI assistance and instructor-performer interaction.
NymeriaMajor egocentric multimodal motion-language dataset from Project Aria.
HOT3DMajor 3D hand-object tracking dataset for AR/VR.
OpenEgo, EgoDex, EgoVerseRecent large-scale egocentric manipulation data for robot learning.
Challenges and Trends in Egocentric Vision: A SurveyRecent comprehensive survey of egocentric tasks, datasets, and open problems.
Bridging Perspectives: Cross-view Collaborative Intelligence with Egocentric-Exocentric VisionSurvey of ego-exo collaboration and paired-capture research.
World Action ModelsSystematizes WAMs across VLA, world models, portable human demonstrations, simulation, and internet-scale egocentric video.

Workshops and Challenges

Recurring venues and challenge hubs where new egocentric tasks and leaderboards appear.

Event / hubFocus
SLT 2026 SmartGlasses ChallengeTwo-track evaluation of timestamped speaker-attributed ASR and spoken-language understanding over 106.98 hours and 714 four-channel smart-glasses audio sessions; data is distributed to registered participants.
EgoLink 2026ACM MM 2026 grand challenge for egocentric social reasoning and interactive tool-using agents, with public labels, code, leaderboard, 4,030 MCQs, and 1,055 evaluation videos.
Ego4D ChallengesEpisodic memory, forecasting, hand-object, social/audio, and other Ego4D tasks.
Ego-Exo4D ChallengesCross-view, correspondence, pose, skill, and proficiency tasks.
EPIC-KITCHENS ChallengesAction recognition, action detection, anticipation, retrieval, and domain adaptation.
HOT3D Challenge3D hand-object tracking and AR/VR pose estimation.
Project AriaAria datasets, tools, and AR sensing ecosystem.
Joint Egocentric Vision (EgoVis) WorkshopCross-dataset egocentric vision workshop (CVPR) spanning Ego4D, Ego-Exo4D, EPIC-KITCHENS, HoloAssist, and more.
CVPR / ICCV / ECCV egocentric and embodied AI workshopsSearch annually for "egocentric", "embodied", "ego-exo", "wearable", and "first-person".

Watchlist

These entries are promising but should be rechecked before treating them as stable public datasets in a paper or benchmark.

ResourceWhy to trackCurrent note
EgoLive2026-04arXiv
WiYHLarge manipulation corpus with rich hand/wrist/depth signalsRelease and license need verification.
OpenEgo2025-09arXiv
EgoVerse2026-04arXiv
EgoDex2025-05ICLR 2026
UMI family: FastUMI / MV-UMI / UMIGen / YUBI / Hoi!Rapidly expanding handheld and wrist-view manipulation interfacesTrack public code, HF/GitHub datasets, licenses, and whether claimed releases are complete.
EgoIntrospect2026-05arXiv
EgoBench2026-05arXiv
EgoAERO / EgoEngine / HumanEgo / EgoGuide / Ego-PiConversion of egocentric videos into robot demonstrations and policiesTrack code, dataset artifacts, and robot-transfer evaluation protocols.
EgoEMG / EgoEVHands / TouchMoment / EgoFun3D / EgoTactile / EgoPressDiff / EgoForce2026 hand/contact/event/3D resourcesTrack GitHub/HF data release, pressure/contact assets, and license.
UCS-Bench / StreamMemBench / ReFocus / Plan Watch RecoverStreaming memory, spatial reasoning, and proactive assistanceTrack full data/code releases and raw-video dependencies.
EgoProx / BARISTA / TAVIS / EgoPoint-Bench / Ego-METASNew 2026 evaluation suites with strong task definitionsTrack stable leaderboards, splits, and license terms.
Causal-Plan-1M / HowToDIV / EgoThink-family / EgoCoT / NoRA / Ego2Web / Pause and ThinkReasoning, planning, safety, and agentic egocentric benchmarksConfirm URLs, licenses, and raw-video dependencies.

Inclusion Rules

  1. 1.Prefer official project pages, GitHub repositories, Hugging Face datasets, arXiv pages, or university dataset pages.
  2. 2.Keep the public status conservative. If raw video is not clearly available, mark partial or watch.
  3. 3.Separate raw datasets from derived annotations and benchmarks.
  4. 4.Record modality, scale, task, license/access friction, and raw-video dependency when known.
  5. 5.Keep the main atlas strictly egocentric: human or animal first-person capture, or ego-exo resources where the ego view is central. List related but non-egocentric resources (robot-only datasets, multi-view robotic benchmarks, autonomous-driving data, general long-video reasoning) under Adjacent and Related Resources.
  6. 6.Keep pure dashcam/autonomous-driving and robot-embodiment data out of the main tables; surface it as context in Adjacent and Related Resources only when it is genuinely useful to egocentric work.
  7. 7.Keep figures text-light. Exact resource names and labels should remain in Markdown/SVG so readers can inspect them directly.

Contributing

Contributions are welcome through pull requests and issues. Please include an official source, a concise description of what the resource contributes, and any known access or license notes.

See `CONTRIBUTING.md` for the inclusion policy and style. Automated checks run in CI to keep the README, catalog data, figures, and public exports consistent; contributors do not need to run local scripts before opening a PR.

Cite This Atlas

If this atlas helps your research or project, a citation or a link back is appreciated. Machine-readable metadata lives in `CITATION.cff`.

bibtex
@misc{he_awesome_egocentric_atlas,
  author       = {He, Chaoyue},
  title        = {Awesome Egocentric Atlas: Egocentric AI Datasets, Benchmarks, Models, and Tools},
  year         = {2026},
  howpublished = {\url{https://github.com/ChaoYue0307/awesome-egocentric-atlas}}
}