opencsg/autohub-benchmark
autohub-benchmark This project designs common use scenarios for web-based code, model, and dataset hosting platforms, and provides corresponding prompts and ground truth. These resources can be used to evaluate the localization performance of visual language models (VLMs) in specialized scenarios. Model Hosting Platform GUI Inference Model Platform Accuracy (%) Error (%) Invalid (%) Completion Rate (%) AriaUI Huggingface 70.8 12.5 6.7 100.0… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/autohub-benchmark.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face