CoolFace
20 results

RL

rl-llm-wiki /knowledge-base RL-for-LLMs Wiki An expert-level, citation-backed knowledge base on reinforcement learning for large language models — RLHF, DPO and offline preference optimization, reward modeling, RLVR and reasoning, training systems, and the failure modes — built collaboratively by autonomous agents. Each topic article is a deep dive written so you can learn the topic from it without reading the underlying papers, with every non-obvious claim cited to a source. Every change lands through a… See the full description on the dataset page: https://huggingface.co/datasets/rl-llm-wiki/knowledge-base.17 likes87k downloads2mo agoHugging Facehuggingface-deep-rl-course /course-images0 likes78k downloads2y agoHugging Faceshihao1895 /bridge-rlds Dataset Structure These datasets are used for MemoryVLA training. This is the standard setting and can be directly used for other models as well.All data follow the RLDS format from the Bridge dataset. bridge_orig — 60k+ episodes, widowx robot robotics0 likes74k downloads11mo agoHugging FaceAnthropic /hh-rlhf Dataset Card for HH-RLHF Dataset Summary This repository provides access to two different kinds of data: Human preference data about helpfulness and harmlessness from Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. These data are meant to train preference (or reward) models for subsequent RLHF training. These data are not meant for supervised training of dialogue agents. Training dialogue agents on these data is likely… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/hh-rlhf.text100K<n<1M2.1k likes39k downloads3y agoHugging FaceMINT-SJTU /RW-RL-Dataset RW-RL Dataset: Real-World Reinforcement Learning for Robots Human-intervention companion dataset: RW-RL-HIL-Dataset (BodenAI) is a separately hosted release of real-world policy rollouts and human corrections. Download it from its own dataset page. RW-RL Dataset is a real-world robot interaction dataset released by Boden Intelligence, Junpu Innovation Center, and the MINT Lab at Shanghai Jiao Tong University. It is designed for a bottleneck that… See the full description on the dataset page: https://huggingface.co/datasets/MINT-SJTU/RW-RL-Dataset.video100K<n<1M10 likes21k downloads2d agoHugging FaceSpreadsheet-RL /Spreadsheet-RL Spreadsheet-RL Dataset Project Page | Paper | GitHub | Model This dataset contains the training and evaluation data used by Spreadsheet-RL, a reinforcement learning framework for spreadsheet agents that edit Excel workbooks with tools and receive outcome-based rewards from workbook recalculation and answer-range comparison. News 🚀 2026-08-01: Released the Spreadsheet-RL-8B checkpoint, scaling SpreadsheetBench Pass@1 from 15.9% for the base model to 16.7%… See the full description on the dataset page: https://huggingface.co/datasets/Spreadsheet-RL/Spreadsheet-RL.textreinforcement-learning10K<n<100K5 likes17k downloads2mo agoHugging Face