datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Motion-o-MCoT-PLM-motion-keyframes
Motion-o-MCoT (PLM + motion keyframes)
Subset of STGR: STR_plm_rdcap rows with <motion in reasoning_process, plus sharded keyframes under videos/stgr/plm/kfs/.
Train split: 3,168 examples (see export_manifest.json in the repo for exact export stats).
Keyframes: JPEGs are stored under shard subfolders (e.g. videos/stgr/plm/kfs/plm_0150/…) so each directory stays under Hugging Face file-count limits. Each key_frames[].path in the JSON is relative to videos/stgr/plm/kfs/ (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/bishoygaloaa/Motion-o-MCoT-PLM-motion-keyframes.motif-qa
MotifQA
Dataset Summary
MotifQA is a synthetic graph question-answering benchmark focused on detecting graph motifs inside small random graphs.
Each example pairs a textual prompt with an answer sentence, a list of nodes highlighted as the motif (when present), and an explicit graph description(nodes and edges).
In this QA dataset, all graphs are homogenous and undirected.
Subsets cover both yes/no motif detection, motif-type classification (house vs 5-cycle), and… See the full description on the dataset page: https://huggingface.co/datasets/naos-ku/motif-qa.motionatlas-bench
MotionAtlas-Bench v1
MotionAtlas-Bench v1 is a video multiple-choice benchmark for motion and target-entity understanding. This public release contains MCQ records, answer keys, media, and target-object masks needed to reproduce the visual grounding settings.
Resources
Paper: https://arxiv.org/abs/2606.29531
Project page: https://kagura-0001.github.io/projects/MotionAtlas/
Code: https://github.com/Kagura-0001/MotionAtlas
MotionAtlas-Data:… See the full description on the dataset page: https://huggingface.co/datasets/maxLWSv2/motionatlas-bench.motibench
MotiBench
MotiBench is a benchmark for evaluating video generation models under physically grounded and commonsense-driven settings. Each image depicts a moment immediately before a physical event, in which a small, localized action is expected to trigger a larger physical response. All images explicitly capture the pre-event state, in which no visible motion has yet occurred, yet the physical configuration strongly implies an imminent interaction.
Sources and Task… See the full description on the dataset page: https://huggingface.co/datasets/shinying/motibench.Wikipedia-Vision-JA
Dataset Card for Wikipedia-Vision-JA
Dataset description
The Wikipedia-Vision-JA is a Vision Language Model dataset generated from Japanese Wikipedia, containing 1.6M pairs of images, captions, and descriptions.
This dataset itself does not contain raw image data. Instead, an image_url is provided for each item.
Format
Wikipedia_Vision_JA.jsonl contains JSON-formatted rows with the following keys:
key: Unique JSON ID
caption: Short caption for the image… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Wikipedia-Vision-JA.
