CoolFace
20 results

prime-rl

PRIME-RL /Eurus-2-RL-Data Eurus-2-RL-Data Links 📜 Paper 📜 Blog 🤗 PRIME Collection Introduction Eurus-2-RL-Data is a high-quality RL training dataset of mathematics and coding problems with outcome verifiers (LaTeX answers for math and test cases for coding). For math, we source from NuminaMath-CoT. The problems span from Chinese high school mathematics to International Mathematical Olympiad competition questions. For coding, we source from APPS, CodeContests, TACO, and Codeforces.… See the full description on the dataset page: https://huggingface.co/datasets/PRIME-RL/Eurus-2-RL-Data.text100K<n<1M59 likes933 downloads2y agoHugging Facejustus27 /prime-rl-code-alltabular10K<n<100K0 likes209 downloads1y agoHugging FacePRIME-RL /Eurus-2-Rollouttext100K<n<1M2 likes153 downloads2y agoHugging FaceprimerL /real_world_sampleimagen<1K1 likes151 downloads1y agoHugging FacePRIME-RL /Eurus-2-SFT-Data Eurus-2-SFT-Data Links 📜 Paper 📜 Blog 🤗 PRIME Collection Introduction Eurus-2-SFT-Data is an action-centric chain-of-thought reasoning dataset, where the policy model chooses one of 7 actions at each step and stops after executing each action. We list the actions as below: Action Name Description ASSESS Analyze current situation, identify key elements and goals ADVANCE Move forward with reasoning - calculate, conclude, or form hypothesis… See the full description on the dataset page: https://huggingface.co/datasets/PRIME-RL/Eurus-2-SFT-Data.text100K<n<1M15 likes144 downloads2y agoHugging FaceprimerL /multimodel_evalimagen<1K0 likes101 downloads1y agoHugging Face