datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
egotempo
EgoTempo
Full-set metadata for lmms-eval task egotempo.
Annotation source: https://raw.githubusercontent.com/google-research-datasets/egotempo/main/egotempo_openQA.json
Raw video location: Ego4D clips (license-gated), clip id in clip_id.
Expected local media root for evaluation: $EGOTEMPO_VIDEO_DIR or $HF_HOME/egotempo.
HLE-Verified
HLE-Verified (HF-native)
This dataset is a Hugging Face-native conversion of skylenage/HLE-Verified at revision becad9f339dfce27df0ebb38e55dabef12ca5735.
Why this exists
The source dataset stores nested verification fields with mixed runtime types (for example 0/1/"uncertain"), which breaks strict Arrow JSON parsing in datasets.load_dataset.
This converted dataset normalizes those fields and publishes split-ready JSONL files for direct use in lmms_eval.
Split… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-eval/HLE-Verified.repcount
RepCount
Full-set metadata for lmms-eval task repcount.
Dataset Contents
This repository contains metadata/annotations only.
Raw video files are not embedded in this repo.
QA Format
Each row is count-based QA for repetition counting:
video / video_id: media reference
question: natural-language counting prompt
count: integer ground-truth repetition count
action_type (optional): action category metadata
In lmms-eval, targets are read from count first… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-eval/repcount.intentqa_lmmsevalvggsound
VGGSound
Full-set metadata for lmms-eval task vggsound.
Dataset Contents
This repository contains annotations and metadata only (YouTube IDs, start seconds, labels, and splits).
Raw audio/video media files are not hosted here.
Raw Media Requirement
To run media-backed evaluation, download original clips from YouTube using metadata fields.
If local media is missing, full evaluation for vggsound cannot run.
Annotation source:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-eval/vggsound.ovr_kinetics
OVR-Kinetics
Full-set metadata for lmms-eval task ovr_kinetics.
Dataset Contents
This repository contains annotations and metadata only (video IDs, start/end times, repetition labels, and splits).
Raw video files are not hosted here.
Raw Media Requirement
To run media-backed evaluation, download original clips from YouTube/Kinetics sources using metadata fields.
If local media is missing, full evaluation for ovr_kinetics cannot run.
Annotation source:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-eval/ovr_kinetics.ssv2
SSv2
Full-set metadata for lmms-eval task ssv2.
Metadata source: https://huggingface.co/datasets/emirgocen/Something-Something-v2
Raw video location: official Something-Something-V2 .webm files (license-gated), not embedded here.
Expected local media root for evaluation: $SSV2_VIDEO_DIR or $HF_HOME/ssv2.
av_asr
AV-ASR
Full-set metadata for lmms-eval task av_asr.
Metadata source: https://huggingface.co/datasets/mattymchen/lrs3-test (LRS3 test mirror)
Raw media location: LRS3 audio/video files (official access is license-gated), media not embedded here.
Expected local media roots for evaluation: $AV_ASR_AUDIO_DIR, $AV_ASR_VIDEO_DIR, or $HF_HOME/av_asr.
