CoolFace
Datasetpublic

kaimeta/wearables_benchmarks

Wearables Benchmarks This page consolidates the GitHub and HuggingFace links to wearables AI benchmarks from Meta. Datasets highly relevant to wearable use cases Benchmark Paper GitHub HuggingFace TLDR WearVox arXiv facebookresearch/wearvox WearVox A benchmark for evaluating voice assistants in realistic wearable scenarios using 3,842 egocentric audio recordings collected via AI glasses across multiple task types like QA, tool calling, and speech… See the full description on the dataset page: https://huggingface.co/datasets/kaimeta/wearables_benchmarks.

sourceHugging Faceunknownupdated 6mo agoView on Hugging Face
1likes24downloads
Dataset Card

Wearables Benchmarks

This page consolidates the GitHub and HuggingFace links to wearables AI benchmarks from Meta.

Datasets highly relevant to wearable use cases

**Benchmark****Paper****GitHub****HuggingFace****TLDR**
WearVoxarXivfacebookresearch/wearvoxWearVoxA benchmark for evaluating voice assistants in realistic wearable scenarios using 3,842 egocentric audio recordings collected via AI glasses across multiple task types like QA, tool calling, and speech translation.
WearVQAarXivWearVQAA benchmark for evaluating visual question answering on wearable devices using 2,520 image-question pairs that reflect egocentric challenges like occlusion, poor lighting, and blur.
CRAGarXivfacebookresearch/CRAGA comprehensive RAG benchmark for factual question answering spanning five domains, eight question categories, and varied entity popularity and temporal dynamics.
CRAG-MMarXivfacebookresearch/CRAG-MMSingle-Turn & Multi-TurnA multimodal conversational benchmark featuring image-based QA across 13 domains with single- and multi-turn conversations captured via smart glasses and public sources.
MemoryQAarXivfacebookresearch/MemoryQAA benchmark for answering recall questions about visual content retrieved from previously stored multimodal memories.
PLM-VideoBencharXivfacebookresearch/perception_modelsA comprehensive video understanding benchmark suite covering fine-grained QA, egocentric smart glasses QA, region captioning, temporal localization, and dense video captioning.

Datasets moderately relevant to wearable use cases

**Benchmark****Paper****GitHub****TLDR**
Head-to-tailarXivfacebookresearch/head-to-tailA benchmark specifically designed to assess how well LLMs incorporate factual knowledge across head, torso, and tail popularity distributions.
VisualLensarXivfacebookresearch/visuallensA recommendation benchmark evaluating systems under task-agnostic visual user histories using data from Google Local and Yelp.
MeetingQAarXivfacebookresearch/AssoMemA benchmark simulating real-world meeting scenarios where multi-turn dialogues form the memory base, paired with diverse QA examples.
SemiBencharXivfacebookresearch/SemiBenchA benchmark for evaluating knowledge extraction quality and its impact on question answering from semi-structured webpages.

License: Please refer to each dataset's license terms and conditions before using these datasets. License information can be found in the respective GitHub repositories and HuggingFace dataset pages.