CoolFace
Datasetpublic

mipal/AVATAR

AVATAR: What’s Making That Sound Right Now? Video-centric Audio-Visual Localization AVATAR stands for Audio-Visual localizAtion benchmark for a spatio-TemporAl peRspective in video. AVATAR is a benchmark dataset designed to evaluate video-centric audio-visual localization (AVL) in complex and dynamic real-world scenarios.Unlike previous benchmarks that rely on static image-level annotations and assume simplified conditions, AVATAR offers high-resolution temporal annotations over… See the full description on the dataset page: https://huggingface.co/datasets/mipal/AVATAR.

sourceHugging Facecc-by-4.0updated 11mo agoView on Hugging Face
1likes72downloads
settings

This repository belongs to mipal on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameAVATAR
visibilitypublic
licencecc-by-4.0
gatedno
ownermipal
Account settings
mipal/AVATAR · CoolFace