petkopetkov/longhorizon-orchestrator-benchmark
LongHorizon Orchestrator Benchmark An offline benchmark for the judgements a manipulation orchestrator delegates to a vision-language model. An orchestrator wraps a frozen low-level policy and replaces a compound instruction ("put everything in the bin") with a stream of single-object subtasks; to do so it must plan (decompose the instruction into subtasks), verify (judge from pixels whether the current subtask is finished), track state (know which goals are already done), and —… See the full description on the dataset page: https://huggingface.co/datasets/petkopetkov/longhorizon-orchestrator-benchmark.
0308
