Pixmind-io/video-to-prompt-benchmark
PixMind Video-to-Prompt Benchmark A small, extensible benchmark schema for evaluating video understanding and prompt reconstruction across advertising, cinematic, product, and social-video tasks. Each record describes the expected analysis structure rather than redistributing third-party media. Contributors should only attach media they own, generated themselves, or can legally redistribute. Fields video_id, video_type, duration_seconds, and shot_count… See the full description on the dataset page: https://huggingface.co/datasets/Pixmind-io/video-to-prompt-benchmark.
PixMind Video-to-Prompt Benchmark
A small, extensible benchmark schema for evaluating video understanding and prompt reconstruction across advertising, cinematic, product, and social-video tasks.
Each record describes the expected analysis structure rather than redistributing third-party media. Contributors should only attach media they own, generated themselves, or can legally redistribute.
Fields
video_id,video_type,duration_seconds, andshot_countvisual_summary,camera_motion, andlighting_and_colorstoryboardgeneric_promptmodel_promptssourceandlicense
Use the accompanying PixMind Video-to-Prompt Lab or run the full PixMind workflow.
Evaluation rubric
Recommended submission format
Contributions should include the structured record, a source and license declaration, and—when redistribution is allowed—a media file or stable URL. Do not submit downloaded social videos without explicit redistribution rights.
For each sample, compare at least one generated reconstruction against the source intent. Record human evaluation separately from automated similarity metrics so that aesthetic preference is not presented as objective accuracy.
Benchmark roadmap
- Expand from synthetic specifications to owned and redistributable reference videos.
- Add product advertising, food, fashion, interior, cinematic, social, and motion-design subsets.
- Publish model-specific prompt adapters for Veo, Seedance, Wan, Kling, and PixVerse.
- Add human evaluation annotations and reconstruction outputs.
- Version benchmark releases so results remain comparable over time.
