yali30/findingdory
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents Karmesh Yadav*, Yusuf Ali*, Gunshi Gupta, Yarin Gal, Zsolt Kira Current vision-language models (VLMs) struggle with long-term memory in embodied tasks. To address this, we introduce FindingDory, a benchmark in Habitat that evaluates memory-based reasoning across 60 long-horizon tasks. In this repo, we release the FindingDory Video Dataset. Each video contains images collected from a… See the full description on the dataset page: https://huggingface.co/datasets/yali30/findingdory.
<center> <a href="https://arxiv.org/abs/2506.15635" target="blank"> <img alt="arXiv" src="https://img.shields.io/badge/arXiv-FindingDory-red?logo=arxiv" height="20" /> </a> <a href="https://findingdory-benchmark.github.io/" target="blank"> <img alt="Website" src="https://img.shields.io/badge/🌎Website-FindingDory-blue.svg" height="20" /> </a> <a href="https://github.com/findingdory-benchmark/findingdory-trl" target="blank"> <img alt="GitHub Code" src="https://img.shields.io/badge/Code-FindingDory--TRL-white?&logo=github&logoColor=white" /> </a> <a href="https://huggingface.co/yali30/findingdory-qwen2.5-VL-3B-finetuned" target="_blank""> <img alt="Huggingface Model" src="https://img.shields.io/badge/Model-FindingDory-yellow?logo=huggingface" /> </a> </center>
<center><h1>FindingDory: A Benchmark to Evaluate Memory in Embodied Agents</h1> <a href="https://www.karmeshyadav.com/">Karmesh Yadav</a>, <a href="https://yusufali98.github.io/">Yusuf Ali</a>, <a href="https://gunshigupta.netlify.app/">Gunshi Gupta</a>, <a href="https://www.cs.ox.ac.uk/people/yarin.gal/website/">Yarin Gal</a>, <a href="https://faculty.cc.gatech.edu/~zk15/">Zsolt Kira</a> </center>
Current vision-language models (VLMs) struggle with long-term memory in embodied tasks. To address this, we introduce FindingDory, a benchmark in Habitat that evaluates memory-based reasoning across 60 long-horizon tasks.
In this repo, we release the FindingDory Video Dataset. Each video contains images collected from a robot’s egocentric view as it navigates realistic indoor environments and interacts with objects. This dataset was used to train and evaluate the high-level agent SFT agent in the FindingDory benchmark.
Usage
from datasets import load_dataset
dataset = load_dataset("yali30/findingdory")Dataset Structure
Notes:
- The validation split contains 60 tasks . The training split only contains 55 task because the 5 “Object Attributes” tasks are withheld from the training set.
- A subsampled version of the dataset (96 frames per episode) is available here.
📄 Citation
@article{yadav2025findingdory,
title = {FindingDory: A Benchmark to Evaluate Memory in Embodied Agents},
author = {Yadav, Karmesh and Ali, Yusuf and Gupta, Gunshi and Gal, Yarin and Kira, Zsolt},
journal = {arXiv preprint arXiv:2506.15635},
year = {2025}
}