Randallhy/RefineCut-Bench
RefineCut-Bench A planning-level benchmark for executable video-editing planning (EMNLP 2026, Plans You Can Check: Verifier-Grounded Learning of an Open-Weight Planner for Executable Video-Editing). A task gives a planner a brief, a real clip pool with schema-constrained captions and metadata, optional music metadata with beat tracks, the current timeline state, and an explicit constraint ledger; the planner emits a RefinePatch (RFC 6902-style JSON Patch over a typed timeline)… See the full description on the dataset page: https://huggingface.co/datasets/Randallhy/RefineCut-Bench.
RefineCut-Bench
A planning-level benchmark for executable video-editing planning (EMNLP 2026, Plans You Can Check: Verifier-Grounded Learning of an Open-Weight Planner for Executable Video-Editing). A task gives a planner a brief, a real clip pool with schema-constrained captions and metadata, optional music metadata with beat tracks, the current timeline state, and an explicit constraint ledger; the planner emits a RefinePatch (RFC 6902-style JSON Patch over a typed timeline) and a deterministic verifier applies it and recomputes every ledger entry. Paper: https://arxiv.org/abs/2608.25622 (EMNLP 2026 Main). Evaluation code and prompt: https://github.com/Lancelot-wy/RefineCut.
Layout
tasks/ all_tasks.jsonl, train.jsonl, dev.jsonl, test.jsonl, split_manifest.json, split_audit.json
eval/ common100_items.jsonl evaluation-ready Common-100 items (ledger, clip-pool metadata, failed
initial state, violated constraints) consumed by the evaluation harness
dev100_items.jsonl checkpoint-selection set
test_items.jsonl all 591 test items in the same format
canonical_clean_ids.json the 92 Common-100 tasks whose canonical id never appears in training
clips/ captions.jsonl (7,971 clips: subject / action / scene / camera / scene_category / motion_intensity /
caption_short / duration / source), clip_alias_maps/, clips_meta/<source>/ (public source identifiers)
music/ music_features.json (BPM, beat times, energy, duration for 499 tracks; no audio)
trajectories/ normalized/<teacher>.jsonl canonicalized multi-teacher trajectories (up to 3 steps x 4 branches)
replayed/<teacher>_replay.jsonl the same trajectories with per-branch verifier replay scores
schemas/ constraint_ledger, editplan, refinepatch, timeline_ir, verifier_output JSON schemas
docs/ metric definitions and the VES formulaTask record
{"task_id": "...", "task_type": "A", "task_subtype": "themed_montage", "brief": "...",
"target_duration": 24.0,
"constraint_ledger": [{"item_id": "...", "type": "must_keep_clip", "spec": {...}, "satisfied": false, "evidence": null}, ...],
"clip_pool": ["clip_0001", ...], "structure": {...}, "canonical_id": "..."}Constraint types: must_keep_clip, must_exclude_clip, target_duration, duration_tolerance, pacing, transition_style, music_sync_bpm, music_sync_beat, must_open_with, must_close_with, no_repeat_within_seconds, max_repeats_per_clip, tag_inclusion, tag_exclusion. An entry is a hard constraint when its spec admits a deterministic pass/fail test; softer entries earn graded credit through the fractional constraint-satisfaction rate.
Evaluation protocol
Closed loop with at most T=3 repair steps, greedy decoding, one fixed PatchPlanner prompt, and the frozen Common-100 list. VES = 0.30 FinalCSR + 0.15 HardPass + 0.15 PASR + 0.15 ReqClipRecall + 0.10 DurationPass + 0.10 TimelineValidity + 0.05 NoRegression (docs/METRIC_DEFINITIONS.md). VES is a protocol-specific executable-planning score; rendered video quality is evaluated separately.
Sources and license
Clips: Pexels (3,219), Panda-70M sample (3,990), Pixabay (255), OpenVid (100), and 5-30 s segments of non-game long videos (407); music features from FMA tracks. Raw video and audio are not redistributed; clips/clips_meta/ gives the public source identifiers. Benchmark metadata, captions, ledgers, and trajectories: CC BY-NC 4.0; sources retain their original licenses.
Citation
@inproceedings{refinecut2026,
title = {Plans You Can Check: Verifier-Grounded Learning of an Open-Weight Planner for Executable Video-Editing},
author = {Wang, Haoyu and Feng, Cheng and Bian, Liuyang and Huang, Ruiyang and Wei, Lei and Wen, Yafei and Chen, Xiaoxin and Tang, Xiaoying},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
year = {2026},
note = {to appear}
}