aippolit/RENT-Qwen-7B
116

RENT Model Details
A model trained using RENT: Reinforcement Learning via Entropy Minimization - an unsupervised RL method that requires no external rewards or ground-truth labels. See our github repo and paper for more info on how this model was trained.
The base model used is Qwen2.5-7B-Instruct.
This model was trained using the aime dataset.
When evaluating this model and the base model on AIME (64 runs on each model), we achieve the following results:
(Note that we report the mean and stderr of the 64 scores the model achieves on AIME)
This checkpoint has not been trained, evaluated, or tested on any other dataset.
