datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepScaleR_Difficulty
Difficulty Estimation on DeepScaleR
We annotate the entire DeepScaleR dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation.
DeepScaleR is a curated dataset of 40,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models.
Difficulty Scoring Method
Difficulty scores are estimated using the… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/DeepScaleR_Difficulty.GSM8K_Difficulty
Difficulty Estimation on DeepScaleR
We annotate the entire GSM8K dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation.
GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/GSM8K_Difficulty.safer-instruct
Safer-Instruct: Aligning Language Models with Automated Preference Data
This repository contains the dataset for the paper titled "Safer-Instruct: Aligning Language Models with Automated Preference Data". Check out our project website here!
Abstract
Reinforcement learning from human feedback (RLHF) is a vital strategy for enhancing model capability in language models. However, annotating preference data for RLHF is a resource-intensive and creativity-demanding process… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/safer-instruct.orz_math_difficulty
Difficulty Estimation on Open Reasoner Zero
We annotate the entire Open Reasoner Zero dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction.
Open Reasoner Zero is a curated a dataset of 57,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models.
Difficulty Scoring Method
Difficulty scores are estimated using… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/orz_math_difficulty.MATH_Difficulty
Difficulty Estimation on MATH
We annotate the entire MATH dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation.
The Mathematics Aptitude Test of Heuristics (MATH) dataset consists of problems from mathematics competitions, including the AMC 10, AMC 12, AIME, and more. Each problem in MATH has a full step-by-step solution, which can be used to teach models to generate… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/MATH_Difficulty.SQuad_University_of_Limerickbook_datalimekiln
