physics-iq
physicsiq-candidatesPhysics-IQ-Verified
Physics-IQ Verified Dataset
This repository hosts the Physics-IQ Verified benchmark data for evaluating physical understanding in generative video models.
Physics-IQ Verified is derived from the original Physics-IQ benchmark dataset.
Original Physics-IQ
Paper: Do generative video models understand physical principles?
Repository: Code | Dataset in Google Cloud
Physics-IQ Verified (Recommended)
Paper: Physics-IQ Verified
Repository: Code | Dataset: Here in this repo :)
We… See the full description on the dataset page: https://huggingface.co/datasets/Anates-Labs-Research/Physics-IQ-Verified.Atom-Harness-Seedance-2.5-Physics-IQ-Verified
Atom Harness (Seedance 2.5) — Physics-IQ Verified, multiframe (v2v), best-practice prompts (bpp)
Physics-IQ Verified score: 56.6 — 198/198 cases, single sample per case (no best-of-N, no selection, no reranking), one run (seed 7).
This dataset is the evidence bundle for a Physics-IQ Verified leaderboard entry: all 198 generated videos, per-case component scores, and the run summary. Scores were computed with the official evaluator (physiq/run_physics_iq.py, verified ground truth… See the full description on the dataset page: https://huggingface.co/datasets/ssaroya/Atom-Harness-Seedance-2.5-Physics-IQ-Verified.atom-1-physics-iq-verified-run01
Atom 1 — Physics-IQ Verified i2v submission (4 runs)
Moving Atoms · Physics-IQ Verified · Image-to-Video
4-run aggregate: 47.06 ± 1.32 (4 independent runs, seeds 42/1042/2042/3042).
Individual run means: 45.44 / 47.27 / 46.88 / 48.64. System frozen across
runs (LoRA weights, planner labels, prompts byte-identical); only the
diffusion seed varies. Current i2v board leader: MiniMax H3 at 39.8 ± 0.3
(2026-08-24) — Atom 1's own base checkpoint.
Atom 1 is MiniMaxAI/MiniMax-H3 (FL2VA… See the full description on the dataset page: https://huggingface.co/datasets/ssaroya/atom-1-physics-iq-verified-run01.Kandinsky-WM-1.0-Physics-IQ-Verified
Kandinsky-WM 1.0 Physics-IQ Verified Artifacts
This dataset contains the artifacts for the Kandinsky-WM 1.0 General Physics result on Physics-IQ Verified.
Model code: https://github.com/kandinskylab/kandinsky-wm
Checkpoint used for generation: https://huggingface.co/kandinskylab/Kandinsky-WM-1.0-I2V-5s-PH
Benchmark code and real GT data: https://github.com/google-deepmind/physics-iq-benchmark
The dataset contains four generated runs with seeds 32768, 1109, 1137, and 16384.… See the full description on the dataset page: https://huggingface.co/datasets/Messimm/Kandinsky-WM-1.0-Physics-IQ-Verified.physics-iq-verified
Isoquant on Physics-IQ Verified
This dataset contains the generated videos and evaluation records supporting
Isoquant's I2V and V2V results on Physics-IQ Verified.
Track
Input
Prompt
Best-of-N
Runs
Result
I2V
One official switch frame
BPP
1
4
53.6869 ± 0.8578
V2V
Complete three-second conditioning video
BPP
1
4
57.1489 ± 0.7887
The two tracks are separate submissions. Each track contains its completed
submission card, the exact descriptions used for generation… See the full description on the dataset page: https://huggingface.co/datasets/isoquant-labs/physics-iq-verified.
