self-play
chess-selfplaychess-selfplay-dataChess-Selfplay2PromptCoT-2.0-SelfPlay-4B-48K
PromptCoT-2.0-SelfPlay Datasets
This repository hosts the self-play datasets used in PromptCoT 2.0 (Scaling Prompt Synthesis for LLM Reasoning).These datasets were created by applying the PromptCoT 2.0 synthesis framework to generate challenging math and programming problems, and then training models through self-play with Direct Preference Optimization (DPO).
PromptCoT-2.0-SelfPlay-4B-48K: 48,113 prompts for Qwen3-4B-Thinking-2507 self-play.
PromptCoT-2.0-SelfPlay-30B-11K: 11… See the full description on the dataset page: https://huggingface.co/datasets/xl-zhao/PromptCoT-2.0-SelfPlay-4B-48K.Reversi-Transformer-1-Selfplay
Reversi-Transformer Self-Play Dataset (m1)
This dataset contains self-play game records generated by Reversi-Transformer-1 playing against itself using MCTS, with C++ bitboard acceleration and multi-process shared-memory batched inference.
It provides 1.8 million board states formatted as TFRecords for training policy and value networks in Reversi AI.
Dataset Summary
Total Samples: ~1,796,729 board positions
Train: 15 TFRecord shards (1,619,475 samples)… See the full description on the dataset page: https://huggingface.co/datasets/rsu/Reversi-Transformer-1-Selfplay.details_azarafrooz__mistral-7b-v2-selfplay-v0
Dataset Card for Evaluation run of azarafrooz/mistral-7b-v2-selfplay-v0
Dataset automatically created during the evaluation run of model azarafrooz/mistral-7b-v2-selfplay-v0 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_azarafrooz__mistral-7b-v2-selfplay-v0.
