26TCA/TCA-Bench
TCA-Bench Official ECCV 2026 benchmark release for the paper "Temporal and Cross-modal Alignment for Enhanced Audiovisual Video Captioning." TCA-Bench is a diagnostic benchmark for audiovisual video captioning. It evaluates base audio/visual perception, audio-visual binding, and cross-modal temporal reasoning using structured ground truth annotations. The benchmark contains 459 anonymized short videos. All annotation files use the anonymized mp4 filename as id, matching files in… See the full description on the dataset page: https://huggingface.co/datasets/26TCA/TCA-Bench.
1173
