CoolFace
Datasetpublic

ianlee1996/pokerbench-rl-dpo

PokerBench RL — Counterfactual DPO Preference Data DPO preference pairs and raw self-play logs for training a Texas Hold'em LLM to exploit non-GTO opponents, addressing the PokerBench paper's Future Work observation that pure SFT models lose to GPT-4-style "donking" strategies. This dataset feeds the ianlee1996/pokerbench-qwen3-14b-lora-dpo checkpoint training. How it was built Self-play (5000 hands): ianlee1996/pokerbench-qwen3-14b-lora-mixed (Qwen3-14B + LoRA… See the full description on the dataset page: https://huggingface.co/datasets/ianlee1996/pokerbench-rl-dpo.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes57downloads

ianlee1996/pokerbench-rl-dpo · main · files are served by the source, never re-hosted here