CoolFace
20 results

rewards

rasinmuhammed /verified-sql-rewards Verified SQL Rewards A text-to-SQL corpus where every reward carries a machine-checkable proof that it is correct. Questions, all independently verified 109,306 Databases 1,400 across 7 schema families Tables / data rows 4,400 / ~19.6 million Unique (question, answer) pairs 102,764 Candidates refused and published 12,150 Verification pass rate 90.00% Trivial baseline (always answer 0) 1.83% Each item is a natural-language question, a gold SQL query… See the full description on the dataset page: https://huggingface.co/datasets/rasinmuhammed/verified-sql-rewards.texttable-question-answering100K<n<1M0 likes291 downloads20d agoHugging FaceShiki258 /hh-rlhf_generation_rewards_iter2text10K<n<100K0 likes175 downloads1y agoHugging Facedebajyotidasgupta /repro-contextual-rollout-bandits-for-reinforcement-learning-with-verifiable-rewards-artifacts Reproduction: Contextual Rollout Bandits for RLVR (ICML 2026, #985) Independent reproduction of "Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards" (Lu, Wang, Chai, Yin, Lin, Chen, Luo, Zhuang, Ban, Wang) — OpenReview weMYE1B16x, arXiv 2602.08499. Part of the Hugging Face × AlphaXiv ICML-2026 reproduction challenge. Official code: github.com/lxd99/CBS_public (verl 0.5.x fork). What CBS is The paper reframes rollout scheduling in RLVR as… See the full description on the dataset page: https://huggingface.co/datasets/debajyotidasgupta/repro-contextual-rollout-bandits-for-reinforcement-learning-with-verifiable-rewards-artifacts.0 likes104 downloads2mo agoHugging Faceeagle0504 /multireward-grpo-gsm8k-rewards-qwen2.5-7b Multi-Reward GRPO — GSM8K Rewards (Qwen2.5-7B-Instruct) Raw rollout-level reward observations from the empirical Section of "Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis". This is the data that produced the headline Theorem 3 (correlation-dependent MSE floor) and Proposition 4 (sign-changing conditioning bias) figures on real LLM rollouts. Each rollout was sampled from Qwen/Qwen2.5-7B-Instruct on GSM8K test prompts at temperature 0.7. What's in… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b.tabulartext-generation10K<n<100K0 likes99 downloads4mo agoHugging Facer2e-edits /deepswe-verifier-merged-with-regression-with-filenames-and-rewards-v2tabular1K<n<10K0 likes98 downloads1y agoHugging FaceDuarteMRAlves /llama-3.1-tulu-3-405b-preference-rewardstext100K<n<1M0 likes79 downloads2y agoHugging Face