datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vitra-ego4d-videoego4d_dataset_2c05c4fcego4d-videoEgoCOT is a large-scale embodied planning dataset, which selected egocentric videos from the Ego4D dataset and corresponding high-quality step-by-step language instructions, which are machine generated, then semantics-based filtered, and finally human-verified.
For mored details, please visit EgoCOT_Dataset.
If you find this dataset useful, please consider citing the paper,
@article{mu2024embodiedgpt,
title={Embodiedgpt: Vision-language pre-training via embodied chain of thought}… See the full description on the dataset page: https://huggingface.co/datasets/wofmanaf/ego4d-video.ego4d_rawego4d-random-views-20k
Ego4D Random Views Dataset
This dataset contains 20,000 random view frames sampled from the Ego4D dataset using a high-performance multi-process generation system.
Dataset Overview
Total Images: 20,000 high-quality frames
Image Format: PNG (1024×1024 resolution)
Source: Ego4D v2 dataset (52,665+ video files)
Sampling Method: Multi-process random sampling with maximum diversity
Generation Time: 797.57 seconds (~13 minutes)
Generation Speed: 25.08 frames/second… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/ego4d-random-views-20k.Ego4D-paxion-addupGST_Ego4Dvitra-ego4d-videoego4d_tfdsEgo4D_audio_subsetbabyllava_v2_instruction_ft_Ego4Dvitra-ego4d-videoego4d_manipulation_v1
ego4d_manipulation_v1 (TsFile)
Apache TsFile version of mderry/ego4d-manipulation-v1.
Records: 25,415
Schema (TsFile structure)
episode_index, task_index (TAG) — device dimension(s).
episode_index (FIELD).
task_index (FIELD).
frame_index (FIELD).
observation_images_ego (FIELD).
observation_state_0 (FIELD).
observation_state_1 (FIELD).
observation_state_2 (FIELD).
observation_state_3 (FIELD).
observation_state_4 (FIELD).
observation_state_5 (FIELD).… See the full description on the dataset page: https://huggingface.co/datasets/THULab/ego4d_manipulation_v1.Ego4Dego4d-manipulation-v1CVPR25-OSGNet-Ego4D-NLQ
Ego4D-NLQ Feature for OSGNet
This repository provides the NLQ features for Object-Shot Enhanced Grounding Network (OSGNet). For video feature and text feature of nlq_v2, we do not redistribute the feature in this repository. Please see nlq_v2/GROUNDNLQ_FEATURE.md for how to obtain it from GroundNLQ.
For installation, dataset preparation, training, and evaluation, please refer to the main repository:
GitHub: https://github.com/iLearn-Lab/CVPR25-OSGNet
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/iLearn-Lab/CVPR25-OSGNet-Ego4D-NLQ.univla-ego4d-rlds-dependenciesLEGO-VLLM-Features-Ego4DThis repo is the VLLM features for egocentric action frame generation on Ego4D dataset.
Please refer to our code on github for instructions on how to use it. More repos are available in this collection.
CVPR25-OSGNet-Ego4D-GoalStep
Ego4D-GoalStep Feature for OSGNet
This repository provides the GoalStep Feature for Object-Shot Enhanced Grounding Network (OSGNet).
For installation, dataset preparation, training, and evaluation, please refer to the main repository:
GitHub: https://github.com/iLearn-Lab/CVPR25-OSGNet
Paper: https://openaccess.thecvf.com/content/CVPR2025/html/Feng_Object-Shot_Enhanced_Grounding_Network_for_Egocentric_Video_CVPR_2025_paper.html
Checkpoints:… See the full description on the dataset page: https://huggingface.co/datasets/iLearn-Lab/CVPR25-OSGNet-Ego4D-GoalStep.ego4dEgo4D_Dataego4d_videomae_L14_feature_fps8
📙 Overview
Ego4d video features extracted by VideoMAE_L14 at 8 fps.
It contains 9645 files, each file (e.g. fffbaeef-577f-45f0-baa9-f10cabf62dfb.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768.
🏋️ How-To-Use
Please refer to code EgoInstructor for details.
🎓 Citation
@article{xu2024retrieval,
title={Retrieval-augmented egocentric video captioning},
author={Xu, Jilan and Huang, Yifei and Hou, Junlin and Chen, Guo… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/ego4d_videomae_L14_feature_fps8.ego4dEGO4D is the world's largest egocentric (first person) video ML dataset and benchmark suite, with 3,600 hrs (and counting) of densely narrated video and a wide range of annotations across five new benchmark tasks. It covers hundreds of scenarios (household, outdoor, workplace, leisure, etc.) of daily life activity captured in-the-wild by 926 unique camera wearers from 74 worldwide locations and 9 different countries. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. The approach to data collection was designed to uphold rigorous privacy and ethics standards with consenting participants and robust de-identification procedures where relevant.videollm-online-chat-ego4d-134k
videollm-online-chat-134k
Introduction
This is the dataset proposed in CVPR 2024 paper: VideoLLM-online: Online Video Large Language Model for Streaming Video. Visit our homepage for paper, demo, code, etc.
The datasets contain 113k streaming narration data and 21k (generated) streaming free-form dialogue data.
The streaming narration data is derived from Ego4D narration, but we cleaned and rephrased them with Llama-3, so that the narration text will not contain strange… See the full description on the dataset page: https://huggingface.co/datasets/chenjoya/videollm-online-chat-ego4d-134k.PACO-ego4dego4d-hand-mano
Dataset summary
Per-frame 3D hand annotations and text captions for 1,302,538 egocentric video clips
drawn from 3,320 Ego4D videos. Every clip is 121 frames at 30 fps (4.03 s) at a
540-pixel short side.
This is the annotation release accompanying Controllable Egocentric Video Generation via
Occlusion-Aware Sparse 3D Hand Joints (ECCV 2026). Each clip carries, for both hands and every
frame: 21 3D joints, MANO pose parameters, root rotation and translation, 2D wrist position
and… See the full description on the dataset page: https://huggingface.co/datasets/bochen123/ego4d-hand-mano.ovo-s-bench-ego4d-benchmark
OVO-S-Bench Ego4D
400-question, 87-video video-only common-schema manifest. The source annotations are CC-BY-4.0.
Source Ego4D videos retain their upstream license and are not included here.
The causal input interval is [0,effective_query_time_s]. The raw official query timestamp is
preserved as raw_query_time_s. Three official whole-second timestamps equal ceil(duration_s)
and are clamped only for the effective endpoint: 0.100000s, 0.366667s, and 0.466667s. No future
frames are… See the full description on the dataset page: https://huggingface.co/datasets/kfkas/ovo-s-bench-ego4d-benchmark.devcv_toolbox_ego4dego4d_train_pair_howto100m
📙 Overview
The metadata for Ego4d training set, with paired howto100m video clips. The ego-exo pair is constructed by choosing the ones with shared nouns/verbs.
Each sample represents a short video clip, which consists of
vid: the initial video id.
start_second: the start timestamp of the narration.
end_second: the end timestamp of the narration.
text: the original narration.
noun: a list containing the index of nouns in the Ego4d noun vocabulary.
verb: a list containing the… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/ego4d_train_pair_howto100m.ego4d
