mipal/AVATAR
AVATAR: What’s Making That Sound Right Now? Video-centric Audio-Visual Localization AVATAR stands for Audio-Visual localizAtion benchmark for a spatio-TemporAl peRspective in video. AVATAR is a benchmark dataset designed to evaluate video-centric audio-visual localization (AVL) in complex and dynamic real-world scenarios.Unlike previous benchmarks that rely on static image-level annotations and assume simplified conditions, AVATAR offers high-resolution temporal annotations over… See the full description on the dataset page: https://huggingface.co/datasets/mipal/AVATAR.
Update README.md
Update README.md
Cleaned and re-zipped metadata archive to fix Hugging Face security scan issues
Update README.md
Upload README.md
Delete README.md
Add zipped video and metadata files
initial commit
