CoolFace
Datasetpublic

ianlee1996/pokerbench-8max-reasoning-traces

PokerBench 8-max — teacher-distilled reasoning traces Reasoning traces for 8-max No-Limit Hold'em decisions, distilled from Claude Sonnet 5 on Bedrock in the STaR style, for training small models to reason about poker prices rather than pattern-match to an action. Method The teacher is not told the answer. It reasons freely from the same prompt production sends, and a trace is kept only if its conclusion matches the target label. Telling the teacher the target… See the full description on the dataset page: https://huggingface.co/datasets/ianlee1996/pokerbench-8max-reasoning-traces.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes45downloads

ianlee1996/pokerbench-8max-reasoning-traces · main · files are served by the source, never re-hosted here