CoolFace
Datasetpublic

tonyliao-meta/WearVQA

WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world Scenarios Paper: WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenariosAuthors: Eun Chang*, Zhuangqun Huang*, Yiwei Liao*, Sagar Ravi Bhavsar*, Amogh Param, Tammy Stark, Adel Ahmadyan, Xiao Yang, Jiaqi Wang, Ahsan Abdullah, Giang Nguyen, Akil Iyer, David Hall, Elissa Li, Shane Moon, Nicolas Scheffer, Kirmani Ahmed, Babak Damavandi… See the full description on the dataset page: https://huggingface.co/datasets/tonyliao-meta/WearVQA.

sourceHugging Faceapache-2.0updated 11mo agoView on Hugging Face
0likes40downloads
Dataset Card

WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world Scenarios

Paper: WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios Authors: Eun Chang, Zhuangqun Huang, Yiwei Liao, Sagar Ravi Bhavsar, Amogh Param, Tammy Stark, Adel Ahmadyan, Xiao Yang, Jiaqi Wang, Ahsan Abdullah, Giang Nguyen, Akil Iyer, David Hall, Elissa Li, Shane Moon, Nicolas Scheffer, Kirmani Ahmed, Babak Damavandi, Rakesh Wanga, Anuj Kumar, Rohit Patel, Xin Luna Dong Affiliations: Meta Reality Labs, Meta Conference: NeurIPS 2025 Datasets and Benchmarks Track


📝 Dataset Summary

WearVQA is the first benchmark specifically designed to evaluate the Visual Question Answering (VQA) capabilities of multi-modal AI assistants on wearable devices like smart glasses. Unlike prior benchmarks that focus on high-quality, third-person imagery, WearVQA reflects the unique challenges of egocentric interaction—where visual inputs may be occluded, poorly lit, unzoomed, or blurry, and questions are grounded in realistic wearable use cases.

  • —2,520 image-question-answer triplets
  • —7 diverse image domains (including text-centric and general scenes)
  • —10 cognitive task types (from basic recognition to complex reasoning)
  • —6 common wearables-specific image quality issues

All questions are designed to be answerable using only the visual input and common sense. The dataset is paired with a rigorous LLM-as-a-judge evaluation framework with 96% labeling accuracy.


💡 Supported Tasks and Leaderboards

  • —Visual Question Answering (VQA)
  • —Robustness to Wearable-Specific Image Quality Issues
  • —Reasoning over Egocentric Visual Inputs

A leaderboard is available for comparing open-source and proprietary multi-modal LLMs on this benchmark.


📊 Dataset Structure

Each example in the dataset contains:

  • —image: The egocentric image (RGB)
  • —question: A natural language question grounded in the image
  • —answer: The ground-truth answer
  • —domain: The image domain (e.g., text-centric, general scene)
  • —task_type: The cognitive task type (e.g., recognition, reasoning)
  • —quality_issue: The type of image quality issue (if any)

📚 Citation

If you use this dataset, please cite:

bibtex
@inproceedings{
chang2025wearvqa,
title={Wear{VQA}: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios},
author={Eun Chang and Zhuangqun Huang and Yiwei Liao and Sagar Ravi Bhavsar and Amogh Param and Tammy Stark and Adel Ahmadyan and Xiao Yang and Jiaqi Wang and Ahsan Abdullah and Giang Nguyen and Akil Iyer and David Patrick hall and Elissa Li and Nicolas SCHEFFER and Ahmed Kirmani and Babak Damavandi and Rakesh Wanga and Anuj Kumar and Rohit Patel and Seungwhan Moon and Xin Luna Dong},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track},
year={2025},
url={https://openreview.net/forum?id=s5p9ByKN1j}
}