Babelscape/PDDL2PRM
PDDL2PRM: Planning-Based Step-Level Supervision for Process Reward Models PDDL2PRM is a large-scale dataset for training and evaluating Process Reward Models (PRMs) with fine-grained, step-level supervision derived from symbolic planning problems. Unlike many PRM datasets that rely on human annotation, LLM judges, or final-answer correctness, PDDL2PRM uses Planning Domain Definition Language (PDDL) problems to generate structured reasoning trajectories whose intermediate steps… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/PDDL2PRM.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face