manzoliw/trucobench-sft
TrucoPaulista SFT v2 Dataset This dataset contains 8,204,630 reasoning-augmented instruction turns generated from 100,000 self-play games of a mathematically optimal heuristic agent (HeuristicAgent) playing Truco Paulista. It is designed to fine-tune Large Language Models (LLMs) to master strategic reasoning, bluffing, and decision-making under imperfect information. Dataset Details Game: Truco Paulista (Brazilian card game) Total Turns/Examples: 8,204,630… See the full description on the dataset page: https://huggingface.co/datasets/manzoliw/trucobench-sft.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face