CoolFace
Datasetpublic

maghzal/PathEval

PathEval: A Benchmark for Evaluating Vision-Language Models as Evaluators for Path Planning Overview Despite their promise to perform complex reasoning, large language models (LLMs) have been shown to have limited effectiveness in end-to-end planning. This has inspired an intriguing question: if these models cannot plan well, can they still contribute to the planning framework as a helpful plan evaluator? In this work, we generalize this question to consider LLMs… See the full description on the dataset page: https://huggingface.co/datasets/maghzal/PathEval.

sourceHugging Facemitupdated 2y agoView on Hugging Face
0likes3.1kdownloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
maghzal/PathEval · CoolFace