datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Video_AMME_ci
Video-AMME CI
Video-AMME is a 50-case CI dataset derived from zhaochenyang20/Video_MME_ci.
Each example keeps the Video-MME video and moves the question, answer
choices, and answer-format instruction into a spoken WAV file.
Files
data/test.jsonl: metadata and source Video-MME references.
audios/*.wav: spoken question/options/instruction.
videos/*.mp4: present only when built with --copy-videos.
Generation
TTS model: fishaudio/s2-pro
Max samples requested: 50… See the full description on the dataset page: https://huggingface.co/datasets/zhaochenyang20/Video_AMME_ci.Video_AMME_ci
Video-AMME CI
Video-AMME is a 50-case CI dataset derived from zhaochenyang20/Video_MME_ci.
Each example keeps the Video-MME video and moves the question, answer
choices, and answer-format instruction into a spoken WAV file.
Files
data/test.jsonl: metadata and source Video-MME references.
audios/*.wav: spoken question/options/instruction.
videos/*.mp4: present only when built with --copy-videos.
Generation
TTS model: fishaudio/s2-pro
Max samples requested: 50… See the full description on the dataset page: https://huggingface.co/datasets/Ratish21/Video_AMME_ci.
