CoolFace
Datasetpublic

JWei05/DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k

DeepScaleR Easy/Medium/Hard — Gemma 4 26B-A4B PT This dataset contains 9,900 unique, deduplicated DeepScaleR math questions for reinforcement-learning experiments. Difficulty is defined by how often the pretrained google/gemma-4-26B-A4B teacher solved each question across eight temperature-1 samples under the same rule-based grader used by the RL training pipeline. The Hub dataset has three configurations—easy, medium, and hard—and each configuration has a train split with 3,000… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes284downloads

JWei05/DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k · main · files are served by the source, never re-hosted here