CoolFace
20 results

self-play

mitermix /chess-selfplaytext1M<n<10M7 likes2k downloads3y agoHugging FaceChristophSchuhmann /chess-selfplay-datatext1K<n<10K0 likes598 downloads3y agoHugging FaceChristophSchuhmann /Chess-Selfplay2text1M<n<10M3 likes291 downloads3y agoHugging Facexl-zhao /PromptCoT-2.0-SelfPlay-4B-48K PromptCoT-2.0-SelfPlay Datasets This repository hosts the self-play datasets used in PromptCoT 2.0 (Scaling Prompt Synthesis for LLM Reasoning).These datasets were created by applying the PromptCoT 2.0 synthesis framework to generate challenging math and programming problems, and then training models through self-play with Direct Preference Optimization (DPO). PromptCoT-2.0-SelfPlay-4B-48K: 48,113 prompts for Qwen3-4B-Thinking-2507 self-play. PromptCoT-2.0-SelfPlay-30B-11K: 11… See the full description on the dataset page: https://huggingface.co/datasets/xl-zhao/PromptCoT-2.0-SelfPlay-4B-48K.text10K<n<100K0 likes196 downloads1y agoHugging Facersu /Reversi-Transformer-1-Selfplay Reversi-Transformer Self-Play Dataset (m1) This dataset contains self-play game records generated by Reversi-Transformer-1 playing against itself using MCTS, with C++ bitboard acceleration and multi-process shared-memory batched inference. It provides 1.8 million board states formatted as TFRecords for training policy and value networks in Reversi AI. Dataset Summary Total Samples: ~1,796,729 board positions Train: 15 TFRecord shards (1,619,475 samples)… See the full description on the dataset page: https://huggingface.co/datasets/rsu/Reversi-Transformer-1-Selfplay.tabularreinforcement-learningn<1K0 likes188 downloads21d agoHugging Faceopen-llm-leaderboard-old /details_azarafrooz__mistral-7b-v2-selfplay-v0 Dataset Card for Evaluation run of azarafrooz/mistral-7b-v2-selfplay-v0 Dataset automatically created during the evaluation run of model azarafrooz/mistral-7b-v2-selfplay-v0 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_azarafrooz__mistral-7b-v2-selfplay-v0.0 likes152 downloads2y agoHugging Face