amd/TTT-Bench
TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games π Paper | π€ Dataset | π Website We introduce TTT-Bench, a new benchmark specifically created to evaluate the reasoning capability of LRMs through a suite of simple and novel two-player Tic-Tac-Toe-style games. Although trivial for humans, these games require basic strategic reasoning, including predicting an opponent's intentions and understanding spatial configurations.β¦ See the full description on the dataset page: https://huggingface.co/datasets/amd/TTT-Bench.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone elseβs repository from here would need an authorised integration and the account holderβs consent, so the link goes to the source instead.
Open discussions on Hugging Face