CoolFace
Datasetpublic

plnguyen2908/AudioVisual-Benchmark-Evaluation

AudioVisual Benchmark Evaluation — evaluation subsets Item-id lists for the audio-visual benchmark subsets used in our reported evaluation tables. Layout <benchmark>/eval_subset.csv item ids evaluated in the paper <benchmark>/media_index.csv id -> media filename(s) <benchmark>/media/ the media files those ids refer to eval_subset.csv holds a single id column keyed to the source benchmark (question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.

sourceHugging Faceotherupdated 22d agoView on Hugging Face
0likes208downloads
Dataset Card

AudioVisual Benchmark Evaluation — evaluation subsets

Item-id lists for the audio-visual benchmark subsets used in our reported evaluation tables.

Layout

<benchmark>/eval_subset.csv   item ids evaluated in the paper
<benchmark>/media_index.csv   id -> media filename(s)
<benchmark>/media/            the media files those ids refer to

eval_subset.csv holds a single id column keyed to the source benchmark (question_id, idx, or index). media/ contains exactly the files those ids reference and nothing else. Questions and answers are not redistributed — take them from the source benchmark and join on the id.

Item count and file count differ by design: several benchmarks ask multiple questions about one clip (AVHBench 4803 items / 2021 videos), while others need more than one file per item (OmniBench is one image plus one audio each). media_index.csv gives the exact mapping.

Benchmarks

folderbenchmarkeval subset nmedia filesid column
avhbenchAVHBench48034042question_id
cmm_fullCMM23163463idx
avsAV-SpeakerBench19271365question_id
omnibenchOmniBench11182236index
videommeVideo-MME (short+medium)1749600question_id
worldsenseWorldSense31181645index

Evaluation subsets

As mentioned in the paper, we evaluate on the short and medium length subset of all of the benchmarks because we believe that it takes a different approach for long video.

Source benchmarks

benchmarksource
AVHBenchSung-Bin et al., AVHBench
CMMCurse of Multi-Modalities
AV-SpeakerBenchplnguyen2908/AV-SpeakerBench
OmniBenchm-a-p/OmniBench
Video-MMElmms-lab/Video-MME
WorldSensehonglyhly/WorldSense

Refer to each source benchmark for its own license and terms.