CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rl-llm-wiki /knowledge-base RL-for-LLMs Wiki An expert-level, citation-backed knowledge base on reinforcement learning for large language models — RLHF, DPO and offline preference optimization, reward modeling, RLVR and reasoning, training systems, and the failure modes — built collaboratively by autonomous agents. Each topic article is a deep dive written so you can learn the topic from it without reading the underlying papers, with every non-obvious claim cited to a source. Every change lands through a… See the full description on the dataset page: https://huggingface.co/datasets/rl-llm-wiki/knowledge-base.17 likes88k downloads2mo agoHugging Face02huggingface-deep-rl-course /course-images0 likes77k downloads2y agoHugging Face03shihao1895 /bridge-rlds Dataset Structure These datasets are used for MemoryVLA training. This is the standard setting and can be directly used for other models as well.All data follow the RLDS format from the Bridge dataset. bridge_orig — 60k+ episodes, widowx robot robotics0 likes57k downloads11mo agoHugging Face04Anthropic /hh-rlhf Dataset Card for HH-RLHF Dataset Summary This repository provides access to two different kinds of data: Human preference data about helpfulness and harmlessness from Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. These data are meant to train preference (or reward) models for subsequent RLHF training. These data are not meant for supervised training of dialogue agents. Training dialogue agents on these data is likely… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/hh-rlhf.text100K<n<1M2.1k likes38k downloads3y agoHugging Face05MINT-SJTU /RW-RL-Dataset RW-RL Dataset: Real-World Reinforcement Learning for Robots Human-intervention companion dataset: RW-RL-HIL-Dataset (BodenAI) is a separately hosted release of real-world policy rollouts and human corrections. Download it from its own dataset page. RW-RL Dataset is a real-world robot interaction dataset released by Boden Intelligence, Junpu Innovation Center, and the MINT Lab at Shanghai Jiao Tong University. It is designed for a bottleneck that… See the full description on the dataset page: https://huggingface.co/datasets/MINT-SJTU/RW-RL-Dataset.video100K<n<1M10 likes20k downloads19h agoHugging Face06mikasa-robo /mikasa-robo-vla-rlds0 likes17k downloads3mo agoHugging Face07Spreadsheet-RL /Spreadsheet-RL Spreadsheet-RL Dataset Project Page | Paper | GitHub | Model This dataset contains the training and evaluation data used by Spreadsheet-RL, a reinforcement learning framework for spreadsheet agents that edit Excel workbooks with tools and receive outcome-based rewards from workbook recalculation and answer-range comparison. News 🚀 2026-08-01: Released the Spreadsheet-RL-8B checkpoint, scaling SpreadsheetBench Pass@1 from 15.9% for the base model to 16.7%… See the full description on the dataset page: https://huggingface.co/datasets/Spreadsheet-RL/Spreadsheet-RL.textreinforcement-learning10K<n<100K5 likes17k downloads2mo agoHugging Face08openvla /modified_libero_rlds Modified LIBERO RLDS Datasets This repository contains the four modified LIBERO datasets used in the OpenVLA fine-tuning experiments, stored in RLDS data format. See Appendix E in the OpenVLA paper for details about the fine-tuning experiments and specific dataset modifications, and see the OpenVLA GitHub README for instructions on how to run OpenVLA in LIBERO environments. Citation BibTeX: @article{kim24openvla, title={OpenVLA: An Open-Source… See the full description on the dataset page: https://huggingface.co/datasets/openvla/modified_libero_rlds.77 likes17k downloads2y agoHugging Face09RLinf /RECAP-Libero10-Task0-48succ-Data2 likes15k downloads5mo agoHugging Face10timaeus /rl-lm-formality-promptstext10K<n<100K0 likes15k downloads4mo agoHugging Face11timaeus /rl-lm-imdb-promptstext10K<n<100K0 likes15k downloads4mo agoHugging Face12openbmb /UltraData-RL-2609 UltraData-RL-2609 📦 UltraData Collection | 🌐 UltraData | 🤗 MiniCPM5 Series English | 中文 📚 Introduction UltraData-RL-2609 is the L3 refined data for reinforcement learning within UltraData's L0-L4 tiered data management framework. Built for the RL stage of MiniCPM5-2B post-training, it complements UltraData-SFT-2605 with verifiable-reward tasks. It is also the training corpus used by JustRL II (Scaling Small LLMs to 128K Reasoning with a Critic)… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraData-RL-2609.text-generation10K<n<100K158 likes15k downloads15d agoHugging Face13timaeus /rl-lm-toxicity-promptstext10K<n<100K0 likes9.3k downloads4mo agoHugging Face14nlile /NuminaMath-1.5-RL-Verifiable Dataset Card for NuminaMath-1.5-RL-Verifiable Dataset Summary NuminaMath-1.5-RL-Verifiable is a curated subset of the NuminaMath-1.5 dataset, specifically filtered to support reinforcement learning applications requiring verifiable outcomes. This collection consists of 131,063 math word problems from the original dataset that meet strict filtering criteria: all problems have definitive numerical answers, validated problem statements and solutions, and come from… See the full description on the dataset page: https://huggingface.co/datasets/nlile/NuminaMath-1.5-RL-Verifiable.texttext-generation100K<n<1M10 likes8.7k downloads1y agoHugging Face15RLinf /LIBERO-assets3dn<1K0 likes8.2k downloads2mo agoHugging Face16RLinf /LIBERO-PRO-assets3dn<1K0 likes8k downloads2mo agoHugging Face17IPEC-COMMUNITY /OpenFly-rlds3 likes7.6k downloads1y agoHugging Face18lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909-completion Matched no-conftest RLVR study 20260909-completion Lossless research records, grouped by model and trajectory type. Only the listed configurations have published records. Canary diagnostics are excluded from study estimates; run status in provenance distinguishes retired diagnostics from active or completed training. Valid failures, refusals and truncations are retained. The train split name is a dataset-loader convention; record_type identifies whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.texttext-generation10K<n<100K1 likes7.2k downloads9d agoHugging Face19czxlovesu03 /rlds_delta0 likes6.9k downloads8mo agoHugging Face20RL-MIND /XHRBench XHRBench Ultra-High-Resolution Remote Sensing Understanding and Reasoning 🤗 Hugging Face · 🤖 ModelScope · 📄 Paper · 💻 Code English | 中文 📚 Introduction XHRBench evaluates fine-grained perception and complex reasoning in multimodal large language models using ultra-high-resolution remote-sensing imagery. This repository retains the name XHRBench and belongs to the same RSHR benchmark project as RSHR-Bench, with a… See the full description on the dataset page: https://huggingface.co/datasets/RL-MIND/XHRBench.imageimage-text-to-text1K<n<10K7 likes6.6k downloads5d agoHugging Face21gavinlaw /rl-run-archive-2026 RL run archive 2026 Archived raw run artifacts (rollout trajectories, rendered frames, policy and optimizer checkpoints, configs, logs) from simulation reinforcement-learning experiments, published for long-term preservation and reproducibility. Layout mirrors the verified backup trees they were copied from: tilde/20260915-102000/ and taurus/20260915-085631/: batched tar archives. Every archive carries a per-file SHA-256 manifest inside it; the batch inventories (9998.json.gz… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/rl-run-archive-2026.tabularn<1K0 likes6.5k downloads13h agoHugging Face22CopyleftCultivars /Agriculture-Agent-RL-Training-Data Agriculture Agent RL Training Data A growing dataset of RL rollout trajectories for LLM agents on natural/regenerative farming — the first RL/trajectory-shaped dataset in the Copyleft Cultivars collection (every prior dataset here is SFT/conversational Q&A). Agents call real tools (primarily cultivars-mcp, a plant-genomics MCP server) across 9 knowledge categories (plus a 10th, organic_chemistry_soil_science, added 2026-08-11, and an 11th, organic_chemistry_synthesis, added… See the full description on the dataset page: https://huggingface.co/datasets/CopyleftCultivars/Agriculture-Agent-RL-Training-Data.text-generation1 likes6.2k downloads24d agoHugging Face23AffineFoundation /rl-pythontext10K<n<100K2 likes5.9k downloads10mo agoHugging Face24RLinf /RPent-memory RPent Memory Memory dataset used by RPent. Memory is organized per robot with a shared contract: <robot>/ ├── MEMORY.md ├── global/ ├── suite/ └── task_only/ ├── <cell>.json ├── <cell>_recipe.jsonl └── <task_key>.md Each robot provides the layers it uses. RoboCasa requires its global file when running with the default task-global policy. Current layout: libero/ ├── MEMORY.md ├── global/ ├── suite/ ├── task_only/ └── task_card/ robocasa/ ├── task_only/ └── global/… See the full description on the dataset page: https://huggingface.co/datasets/RLinf/RPent-memory.robotics2 likes5.9k downloads4d agoHugging Face25PrimeIntellect /Reverse-Text-RL Reverse-Text-RL A small, scrappy RL dataset used in prime-rl's CI to debug RL training asking a model to reverse small sentences character-by-character. Follows the general format of PrimeIntellect/Reverse-Text-SFT The following script was used to generate the dataset. from datasets import Dataset, load_dataset dataset = load_dataset("willcb/R1-reverse-wikipedia-paragraphs-v1-1000", split="train") prompt = "Reverse the text character-by-character. Put your answer in… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Reverse-Text-RL.textquestion-answering1K<n<10K2 likes5.6k downloads1y agoHugging Face26SynthLabsAI /Big-Math-RL-Verifiedgated Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models Big-Math is the largest open-source dataset of high-quality mathematical problems, curated specifically for reinforcement learning (RL) training in language models. With over 250,000 rigorously filtered and verified problems, Big-Math bridges the gap between quality and quantity, establishing a robust foundation for advancing reasoning in LLMs. Request Early Access to Private… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/Big-Math-RL-Verified.textquestion-answering100K<n<1M243 likes5.5k downloads1y agoHugging Face27SaifPunjwani /slo-rlvr-results0 likes5.4k downloads6m agoHugging Face28PrimeIntellect /Multi-SWE-RL-Verified Multi-SWE-RL-Verified Gold-patch-validated subset of PrimeIntellect/Multi-SWE-RL-Reupload (ByteDance's Multi-SWE-RL): 2,232 / 4,703 rows across C, Go, Java, JavaScript, Rust, and TypeScript that produce a clean reward signal end-to-end. Default dataset of the multiswe_v1 taskset. Changes vs upstream Starting from the 4,703-row re-upload: C++ dropped wholesale — 0/449 rows passed gold-patch validation in pass 1; the images are broken for scoring, not merely… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Multi-SWE-RL-Verified.tabulartext-generation1K<n<10K4 likes5.3k downloads3mo agoHugging Face29rl-rag /hle-gpt-oss-120b-no-python-260222 hle-gpt-oss-120b-no-python-260222 Deep research agent evaluation on rl-rag/hle_text_only (test split). Results Metric Value pass@4 47.9% avg@4 26.6% Trajectory accuracy 26.6% (2292/8632) Questions 2158 Trajectories 8632 (4 per question) Avg tool calls 14.5 Full conversations ❌ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domains huggingface.co Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-no-python-260222.tabular1K<n<10K1 likes4.8k downloads7mo agoHugging Face30KodCode /KodCode-Light-RL-10K 🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning. 🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-Light-RL-10K.tabularquestion-answering10K<n<100K9 likes4.6k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.