plnguyen2908/AudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets Item-id lists for the audio-visual benchmark subsets used in our reported evaluation tables. Layout <benchmark>/eval_subset.csv item ids evaluated in the paper <benchmark>/media_index.csv id -> media filename(s) <benchmark>/media/ the media files those ids refer to eval_subset.csv holds a single id column keyed to the source benchmark (question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer toeval_subset.csv holds a single id column keyed to the source benchmark (question_id, idx, or index). media/ contains exactly the files those ids reference and nothing else. Questions and answers are not redistributed — take them from the source benchmark and join on the id.
Item count and file count differ by design: several benchmarks ask multiple questions about one clip (AVHBench 4803 items / 2021 videos), while others need more than one file per item (OmniBench is one image plus one audio each). media_index.csv gives the exact mapping.
Benchmarks
Evaluation subsets
As mentioned in the paper, we evaluate on the short and medium length subset of all of the benchmarks because we believe that it takes a different approach for long video.
Source benchmarks
Refer to each source benchmark for its own license and terms.
