CoolFace
Datasetpublic

XingYing-stack/TIPS-Training-Data

TIPS Training Data Outcome-labeled training trajectories used by TIPS (Thinking-Induced Process Supervision). Configuration File Examples Math math/train.parquet 3,200 Agent agent/train.parquet 2,905 TIPS trains a generative reward model to produce a reasoning chain, step-level labels, and an outcome label while using only outcome correctness as the reinforcement-learning reward. Usage from datasets import load_dataset math_data =… See the full description on the dataset page: https://huggingface.co/datasets/XingYing-stack/TIPS-Training-Data.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes50downloads

XingYing-stack/TIPS-Training-Data · main · files are served by the source, never re-hosted here