mipal/AVATAR
AVATAR: What’s Making That Sound Right Now? Video-centric Audio-Visual Localization AVATAR stands for Audio-Visual localizAtion benchmark for a spatio-TemporAl peRspective in video. AVATAR is a benchmark dataset designed to evaluate video-centric audio-visual localization (AVL) in complex and dynamic real-world scenarios.Unlike previous benchmarks that rely on static image-level annotations and assume simplified conditions, AVATAR offers high-resolution temporal annotations over… See the full description on the dataset page: https://huggingface.co/datasets/mipal/AVATAR.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face