CoolFace
Datasetpublic

opencsg/autohub-benchmark

autohub-benchmark This project designs common use scenarios for web-based code, model, and dataset hosting platforms, and provides corresponding prompts and ground truth. These resources can be used to evaluate the localization performance of visual language models (VLMs) in specialized scenarios. Model Hosting Platform GUI Inference Model Platform Accuracy (%) Error (%) Invalid (%) Completion Rate (%) AriaUI Huggingface 70.8 12.5 6.7 100.0… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/autohub-benchmark.

sourceHugging Faceupdated 2y agoView on Hugging Face
1likes158downloads
prompt2.json10 linesDownload Raw Back to model_platform_data
1{2  "1_home.png": "Click on the tab to navigate to search models.",3  "2_0_models.png": "Check the pricing bottom",4  "2_1_deepseek-v3.png": "Click the 'copy' button to copy the name of prompt names",5  "2_2_deepseek-v3_modefile.png": "Click the community bottom。",6  "2_5_deepseek-v3_collection.png": "Click the model With the largest number of downloads.",7  "3_0_datasets.png": "Click the model With the largest number of downloads.",8  "3_1_opencsg-chinese-fineweb-edu-v2.png": "Click the train bottom.",9  "3_2_opencsg-smoltalk-chinese.png": "Click the second row of conversations list."10}