CoolFace
Datasetpublic

yali30/findingdory

FindingDory: A Benchmark to Evaluate Memory in Embodied Agents Karmesh Yadav*, Yusuf Ali*, Gunshi Gupta, Yarin Gal, Zsolt Kira Current vision-language models (VLMs) struggle with long-term memory in embodied tasks. To address this, we introduce FindingDory, a benchmark in Habitat that evaluates memory-based reasoning across 60 long-horizon tasks. In this repo, we release the FindingDory Video Dataset. Each video contains images collected from a… See the full description on the dataset page: https://huggingface.co/datasets/yali30/findingdory.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
3likes1.4kdownloads
Dataset Card

<center> <a href="https://arxiv.org/abs/2506.15635" target="blank"> <img alt="arXiv" src="https://img.shields.io/badge/arXiv-FindingDory-red?logo=arxiv" height="20" /> </a> <a href="https://findingdory-benchmark.github.io/" target="blank"> <img alt="Website" src="https://img.shields.io/badge/🌎Website-FindingDory-blue.svg" height="20" /> </a> <a href="https://github.com/findingdory-benchmark/findingdory-trl" target="blank"> <img alt="GitHub Code" src="https://img.shields.io/badge/Code-FindingDory--TRL-white?&logo=github&logoColor=white" /> </a> <a href="https://huggingface.co/yali30/findingdory-qwen2.5-VL-3B-finetuned" target="_blank""> <img alt="Huggingface Model" src="https://img.shields.io/badge/Model-FindingDory-yellow?logo=huggingface" /> </a> </center>

<center><h1>FindingDory: A Benchmark to Evaluate Memory in Embodied Agents</h1> <a href="https://www.karmeshyadav.com/">Karmesh Yadav</a>, <a href="https://yusufali98.github.io/">Yusuf Ali</a>, <a href="https://gunshigupta.netlify.app/">Gunshi Gupta</a>, <a href="https://www.cs.ox.ac.uk/people/yarin.gal/website/">Yarin Gal</a>, <a href="https://faculty.cc.gatech.edu/~zk15/">Zsolt Kira</a> </center>

Current vision-language models (VLMs) struggle with long-term memory in embodied tasks. To address this, we introduce FindingDory, a benchmark in Habitat that evaluates memory-based reasoning across 60 long-horizon tasks.

In this repo, we release the FindingDory Video Dataset. Each video contains images collected from a robot’s egocentric view as it navigates realistic indoor environments and interacts with objects. This dataset was used to train and evaluate the high-level agent SFT agent in the FindingDory benchmark.

Usage

from datasets import load_dataset
dataset = load_dataset("yali30/findingdory")

Dataset Structure

Field nameDescription
ep\_idEpisode id.
videoRelative path of the video clip.
questionQuestion posed to the agent based on the episode.
answerGround-truth answer stored as a list of image indices
task\_idIdentifier indicating which task template the episode belongs to (string).
high\_level\_categoryHigl-task task category label. (Options: Single-Goal Spatial Tasks, Single-Goal Temporal Tasks, Multi-Goal Tasks).
low\_level\_categoryFine-grained task category label. (Example categories: Interaction-Order, Room Visitation, etc)
num\_interactionsNumber of objects the robot interacts with, during the experience collection.

Notes:

  • The validation split contains 60 tasks . The training split only contains 55 task because the 5 “Object Attributes” tasks are withheld from the training set.
  • A subsampled version of the dataset (96 frames per episode) is available here.

📄 Citation

@article{yadav2025findingdory,
  title     = {FindingDory: A Benchmark to Evaluate Memory in Embodied Agents},
  author    = {Yadav, Karmesh and Ali, Yusuf and Gupta, Gunshi and Gal, Yarin and Kira, Zsolt},
  journal   = {arXiv preprint arXiv:2506.15635},
  year      = {2025}
}