CoolFace
Datasetpublic

PTeterwak/action-reward-models-data

Action Reward Models for Web Agents Minimal, self-contained repo for the action reward model (ARM) study: generate per-step candidate-action data from a web-agent policy, train two kinds of reward models on teacher labels, and use them to pick actions at inference time. Everything here was extracted from two production pipelines ("eras") and trimmed to the essential path. Written to be read by an AI assistant picking this up cold — file paths, gotchas, and provenance are spelled… See the full description on the dataset page: https://huggingface.co/datasets/PTeterwak/action-reward-models-data.

sourceHugging Faceupdated 15d agoView on Hugging Face
0likes63downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
PTeterwak/action-reward-models-data · CoolFace