PTeterwak/action-reward-models-data
Action Reward Models for Web Agents Minimal, self-contained repo for the action reward model (ARM) study: generate per-step candidate-action data from a web-agent policy, train two kinds of reward models on teacher labels, and use them to pick actions at inference time. Everything here was extracted from two production pipelines ("eras") and trimmed to the essential path. Written to be read by an AI assistant picking this up cold — file paths, gotchas, and provenance are spelled… See the full description on the dataset page: https://huggingface.co/datasets/PTeterwak/action-reward-models-data.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face