CoolFace
Datasetpublic

Pixmind-io/video-to-prompt-benchmark

PixMind Video-to-Prompt Benchmark A small, extensible benchmark schema for evaluating video understanding and prompt reconstruction across advertising, cinematic, product, and social-video tasks. Each record describes the expected analysis structure rather than redistributing third-party media. Contributors should only attach media they own, generated themselves, or can legally redistribute. Fields video_id, video_type, duration_seconds, and shot_count… See the full description on the dataset page: https://huggingface.co/datasets/Pixmind-io/video-to-prompt-benchmark.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes31downloads
Dataset Card

PixMind Video-to-Prompt Benchmark

A small, extensible benchmark schema for evaluating video understanding and prompt reconstruction across advertising, cinematic, product, and social-video tasks.

Each record describes the expected analysis structure rather than redistributing third-party media. Contributors should only attach media they own, generated themselves, or can legally redistribute.

Fields

  • —video_id, video_type, duration_seconds, and shot_count
  • —visual_summary, camera_motion, and lighting_and_color
  • —storyboard
  • —generic_prompt
  • —model_prompts
  • —source and license

Use the accompanying PixMind Video-to-Prompt Lab or run the full PixMind workflow.

Evaluation rubric

DimensionQuestionScore
Scene coverageDoes the analysis capture every meaningful shot?1–5
Temporal accuracyAre shot boundaries and ordering correct?1–5
Subject fidelityAre subjects, products, and actions correctly described?1–5
Camera languageAre framing and camera movements correctly identified?1–5
Lighting and colorDoes the prompt preserve the visual treatment?1–5
Prompt usabilityCan the result be used with minimal rewriting?1–5
Model adaptationDoes each model-specific prompt reflect that model's prompting style?1–5

Recommended submission format

Contributions should include the structured record, a source and license declaration, and—when redistribution is allowed—a media file or stable URL. Do not submit downloaded social videos without explicit redistribution rights.

For each sample, compare at least one generated reconstruction against the source intent. Record human evaluation separately from automated similarity metrics so that aesthetic preference is not presented as objective accuracy.

Benchmark roadmap

  • —Expand from synthetic specifications to owned and redistributable reference videos.
  • —Add product advertising, food, fashion, interior, cinematic, social, and motion-design subsets.
  • —Publish model-specific prompt adapters for Veo, Seedance, Wan, Kling, and PixVerse.
  • —Add human evaluation annotations and reconstruction outputs.
  • —Version benchmark releases so results remain comparable over time.