polowitty/surfPRM-data
surfPRM-data SFT training data for surfPRM, a process reward model (PRM) for web agents. 9,821 pairwise step-preference samples stored as a single JSON array Each sample has one field conversation with system / user / assistant turns: the user turn contains the task intent, the current page's accessibility tree (AXTree), the previous action trajectory, the start/current URLs, and two candidate next actions; the assistant turn contains structured… See the full description on the dataset page: https://huggingface.co/datasets/polowitty/surfPRM-data.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face