lime-nlp/DeepScaleR_Difficulty
Difficulty Estimation on DeepScaleR We annotate the entire DeepScaleR dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation. DeepScaleR is a curated dataset of 40,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models. Difficulty Scoring Method Difficulty scores are estimated using… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/DeepScaleR_Difficulty.
11137
1version https://git-lfs.github.com/spec/v12oid sha256:55ca1e89504a40a4903df344b989ca5e1c26b28ff1419072baae3962fed104173size 170261808624 