CoolFace
Datasetpublic

WaitHZ/SSPO-data

SSPO Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents 📄 arXiv • 💻 Code • 🤗 Dataset 🌟Overview Deep search agents operate over trajectories spanning dozens of information-seeking steps, but standard reinforcement learning provides only a single outcome reward for the entire trajectory. This sparse signal makes it difficult to determine which intermediate reasoning and tool-use actions should be reinforced… See the full description on the dataset page: https://huggingface.co/datasets/WaitHZ/SSPO-data.

sourceHugging Facemitupdated 9d agoView on Hugging Face
1likes126downloads
settings

This repository belongs to WaitHZ on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameSSPO-data
visibilitypublic
licencemit
gatedno
ownerWaitHZ
Account settings
WaitHZ/SSPO-data · CoolFace