CoolFace
Datasetpublic

ssaroya/Atom-Harness-Seedance-2.5-Physics-IQ-Verified

Atom Harness (Seedance 2.5) — Physics-IQ Verified, multiframe (v2v), best-practice prompts (bpp) Physics-IQ Verified score: 56.6 — 198/198 cases, single sample per case (no best-of-N, no selection, no reranking), one run (seed 7). This dataset is the evidence bundle for a Physics-IQ Verified leaderboard entry: all 198 generated videos, per-case component scores, and the run summary. Scores were computed with the official evaluator (physiq/run_physics_iq.py, verified ground truth… See the full description on the dataset page: https://huggingface.co/datasets/ssaroya/Atom-Harness-Seedance-2.5-Physics-IQ-Verified.

sourceHugging Faceunknownupdated 1mo agoView on Hugging Face
3likes486downloads
Dataset Card

Atom Harness (Seedance 2.5) — Physics-IQ Verified, multiframe (v2v), best-practice prompts (bpp)

Physics-IQ Verified score: 56.6 — 198/198 cases, single sample per case (no best-of-N, no selection, no reranking), one run (seed 7).

This dataset is the evidence bundle for a Physics-IQ Verified leaderboard entry: all 198 generated videos, per-case component scores, and the run summary. Scores were computed with the official evaluator (physiq/run_physics_iq.py, verified ground truth, final_score_view).

For reference, the same run scored with the official evaluator on the original Physics-IQ benchmark (original ground truth, original aggregation) gives 62.6 — evaluator outputs for both variants are included under results/official_evaluator/.

System (declared in full)

This is a composite system — a prompting harness around a hosted video model, in the same family as the leaderboard's best-of-N entries, and it uses an LLM component:

  • —Generator: Seedance 2.5 (dreamina-seedance-2-5-260628 via BytePlus ModelArk), true video-extension task (omni_reference_task_type: "extend"), conditioned on the full 3 s take-1 conditioning videos. 720p, 5 s output (exactly 120 frames @ 24 fps), watermark and audio off, seed 7.
  • —Prompt harness (bpp custom templater): the scene text comes verbatim from the benchmark's descriptions/best_practice/descriptions_base.csv; a Seedance-specific templater re-orders those fields and adds fixed clause templates (an extension trigger, a continuation clause, and a timed clause in the "0s–Xs:" format). A VLM planner (Claude Opus 5) watches only the conditioning clip (a declared input of the v2v track) and the base description fields, and outputs exactly two values per case to fill those templates: a motion class (settling / continuous / quiescent) and a predicted stop time. No scene text is invented; the planner writes no free prose.
  • —Sampling: one generation per case. No candidate selection of any kind.

Integrity

  • —The planner and generator see only the conditioning videos and the official description text — the declared inputs of the v2v track. The testing videos (take 1 and take 2) are never accessed at generation time.
  • —Submitted videos are exactly the 5-second continuation (conditioning segment excluded), 120 frames @ 24 fps, benchmark ID naming preserved.
  • —Known limitations, kept as part of the measured system: the VLM mislabels two rotating scenarios as quiescent; optics/reflection scenarios remain the generator's weakest class.

Contents

videos/atom-harness-seedance-bpp-run_01/   198 generated MP4s (5.000 s each)
results/run_01_per_case.csv               per-case component scores (harness scorer)
results/summary_198.json                  run summary
results/official_evaluator/               official physiq evaluator outputs (CSV + metrics JSON)