CoolFace
Modelpublic

parthh01/chess-llm-tournament

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes17downloads
Model Card

Chess GRPO Trained Model

This model has been trained using Group Relative Policy Optimization (GRPO) to play chess. It was trained to generate chess moves in JSON format with reasoning.

Model Details

  • Model Type: PEFT (merged)
  • Training Method: GRPO (Group Relative Policy Optimization)
  • Task: Chess move generation with evaluation reasoning
  • Source Path: ./grpooutput/skill6-final