CoolFace
Datasetpublic

mipal/AVATAR

AVATAR: What’s Making That Sound Right Now? Video-centric Audio-Visual Localization AVATAR stands for Audio-Visual localizAtion benchmark for a spatio-TemporAl peRspective in video. AVATAR is a benchmark dataset designed to evaluate video-centric audio-visual localization (AVL) in complex and dynamic real-world scenarios.Unlike previous benchmarks that rely on static image-level annotations and assume simplified conditions, AVATAR offers high-resolution temporal annotations over… See the full description on the dataset page: https://huggingface.co/datasets/mipal/AVATAR.

sourceHugging Facecc-by-4.0updated 11mo agoView on Hugging Face
1likes72downloads
8 commits on main
a7f014611mo ago

Update README.md

hahyeon610
b4247f21y ago

Update README.md

hahyeon610
c0ed3fe1y ago

Cleaned and re-zipped metadata archive to fix Hugging Face security scan issues

hahyeon610
ca6d2221y ago

Update README.md

hahyeon610
4f706ee1y ago

Upload README.md

hahyeon610
12efe891y ago

Delete README.md

hahyeon610
f89df011y ago

Add zipped video and metadata files

hahyeon610
ac7d17b1y ago

initial commit

hahyeon610