zhanwenchen/personalgui-accel-data
GUI-agent evaluation: code, per-episode records, and captured media This repository holds everything behind a study of whether a language model can operate real software when its only input is a video feed of the screen and its only output is a keyboard and mouse. The rendered report lives in a separate Space; this is the material it is computed from. What the setup is Two machines, joined by two cables and no software link. One runs the agent. The other runs the… See the full description on the dataset page: https://huggingface.co/datasets/zhanwenchen/personalgui-accel-data.
Merge HuggingFace dataset initialisation; union of LFS patterns
GUI-agent evaluation rig: code, per-episode records, media and findings
initial commit
