Gradygu3u/VSI-Super-Wild-Benchmark
VSI-Super-Wild This repository stores the public video assets and the current lmms-eval benchmark release for Toward Supersensing. Current benchmark release Version: cambw_v2_recheck_20260409_contentfix_mc4 Location: benchmarks/cambw_v2_recheck_20260409_contentfix_mc4/ QA count: 12021 Split count: part1_long: 511 part2_3_short: 11510 Layout videos/: public video files benchmarks/cambw_v2_recheck_20260409_contentfix_mc4/data/: materialized… See the full description on the dataset page: https://huggingface.co/datasets/Gradygu3u/VSI-Super-Wild-Benchmark.
VSI-Super-Wild
This repository stores the public video assets and the current lmms-eval benchmark release for Toward Supersensing.
Current benchmark release
- Version:
cambw_v2_recheck_20260409_contentfix_mc4 - Location:
benchmarks/cambw_v2_recheck_20260409_contentfix_mc4/ - QA count:
12021 - Split count:
part1_long:511part2_3_short:11510
Layout
videos/: public video filesbenchmarks/cambw_v2_recheck_20260409_contentfix_mc4/data/: materialized lmms-eval jsonl filesbenchmarks/cambw_v2_recheck_20260409_contentfix_mc4/source/: source QA json files used to build the benchmarkbenchmarks/cambw_v2_recheck_20260409_contentfix_mc4/manifest.json: benchmark manifest
Code
The corresponding lmms-eval integration and runner scripts are published in:
- GitHub:
Grady10086/Toward-Supersensing
Notes
- This benchmark release is the corrected content-fix version with balanced 4-way multiple-choice ordering.
- Existing model results from older benchmark versions should be backfilled or delta-rerun against this version rather than mixed directly.
