datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pokerbench-rl-dpo
PokerBench RL — Counterfactual DPO Preference Data
DPO preference pairs and raw self-play logs for training a Texas Hold'em LLM
to exploit non-GTO opponents, addressing the PokerBench paper's Future Work observation that pure SFT models lose to GPT-4-style "donking" strategies.
This dataset feeds the ianlee1996/pokerbench-qwen3-14b-lora-dpo checkpoint training.
How it was built
Self-play (5000 hands): ianlee1996/pokerbench-qwen3-14b-lora-mixed (Qwen3-14B + LoRA… See the full description on the dataset page: https://huggingface.co/datasets/ianlee1996/pokerbench-rl-dpo.arena-poker-reasoned-decisions-v0
DevFun Arena Poker - Reasoned Decision Traces (v0)
1000 agent decision traces from live 6-max No-Limit Texas Hold'em on the
dev.fun AI-agent poker Arena. Each row is one agent's decision at one
moment in one hand, paired with the structured rationale the agent emitted for that action.
This is a small curated SAMPLE for researchers to judge whether the full data is useful.
Each decision is enriched with full per-seat table state (every seat's stack at decision time),
all-in… See the full description on the dataset page: https://huggingface.co/datasets/dannyobito/arena-poker-reasoned-decisions-v0.
