WaitHZ/SSPO-data
SSPO Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents 📄 arXiv • 💻 Code • 🤗 Dataset 🌟Overview Deep search agents operate over trajectories spanning dozens of information-seeking steps, but standard reinforcement learning provides only a single outcome reward for the entire trajectory. This sparse signal makes it difficult to determine which intermediate reasoning and tool-use actions should be reinforced… See the full description on the dataset page: https://huggingface.co/datasets/WaitHZ/SSPO-data.
This repository belongs to WaitHZ on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
