datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
physical-ai-bench-generation
Physical AI Bench - Generation
Paper | Code
Dataset Description
The PAI-Bench is a benchmark to measure the progress of world models quantitatively.
The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.ship-dataset
ShipBench: A Drawing-Grounded VLM Benchmark for Ship Structural Reasoning
ShipBench is a metadata-grounded vision-language benchmark on parametrically-generated ship structural drawings. Six commercial ship types × nine drawing-grounded sub-tasks × deterministic ground truth derived directly from the generator's input dictionary (no human annotation, no rule-citation labels).
Quick reference
Total candidates: 6{,}450 across 6 ship types (Tanker, VLCC, BULKC, CNTR, LNGC… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous1383/ship-dataset.ShieldVLMMuSEAgent-EvalJapanese_Mangamotibench
MotiBench
MotiBench is a benchmark for evaluating video generation models under physically grounded and commonsense-driven settings. Each image depicts a moment immediately before a physical event, in which a small, localized action is expected to trigger a larger physical response. All images explicitly capture the pre-event state, in which no visible motion has yet occurred, yet the physical configuration strongly implies an imminent interaction.
Sources and Task… See the full description on the dataset page: https://huggingface.co/datasets/shinying/motibench.SemiFA-930WM-811K images: download via scripts/download_wm811k.py. MixedWM38 images: download via scripts/download_mixedwm38.py. Synthetic images are included in this repository.
Link to github repo to find these scripts-https://github.com/Shivamckaushik/SemiFA
behance_jsonllava_finetuning_dataset_for_text_extractionWebVoyager-Trajectories-Qwen3.5-Omni
WebVoyager + GAIA Agent Trajectories (Qwen3.5-Omni)
Full browser-agent trajectories for 733 tasks (643 WebVoyager + 90 GAIA-web), produced by a browser-use agent driven by qwen3.5-omni-plus-2026-03-15 (multimodal, via Alibaba DashScope), run in a real headed browser. Every step records the exact LLM context (including the screenshot the model saw) and the action taken, plus a reference-grounded success verdict.
Results
Judged by the WebVoyager reference-grounded… See the full description on the dataset page: https://huggingface.co/datasets/shiqihe/WebVoyager-Trajectories-Qwen3.5-Omni.shiny-cards-produceShia_Photography
The Shia Photography Dataset
This dataset features over 6,000 images that capture the spiritual and cultural richness of Shia events and holy shrines, all commissioned for the Holy Shrine of Imam Hussain (A.S.).
With more than 12 photographers contributing to this collection, the images offer a diverse range of perspectives. Each photograph is accompanied by a detailed, descriptive caption to provide objective context.
Caption Generation
The captions were… See the full description on the dataset page: https://huggingface.co/datasets/Koratahi/Shia_Photography.
