CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Anthropic /hh-rlhf Dataset Card for HH-RLHF Dataset Summary This repository provides access to two different kinds of data: Human preference data about helpfulness and harmlessness from Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. These data are meant to train preference (or reward) models for subsequent RLHF training. These data are not meant for supervised training of dialogue agents. Training dialogue agents on these data is likely… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/hh-rlhf.text100K<n<1M2.1k likes39k downloads3y agoHugging Face02Spreadsheet-RL /Spreadsheet-RL Spreadsheet-RL Dataset Project Page | Paper | GitHub | Model This dataset contains the training and evaluation data used by Spreadsheet-RL, a reinforcement learning framework for spreadsheet agents that edit Excel workbooks with tools and receive outcome-based rewards from workbook recalculation and answer-range comparison. News 🚀 2026-08-01: Released the Spreadsheet-RL-8B checkpoint, scaling SpreadsheetBench Pass@1 from 15.9% for the base model to 16.7%… See the full description on the dataset page: https://huggingface.co/datasets/Spreadsheet-RL/Spreadsheet-RL.textreinforcement-learning10K<n<100K5 likes17k downloads2mo agoHugging Face03timaeus /rl-lm-formality-promptstext10K<n<100K0 likes15k downloads4mo agoHugging Face04timaeus /rl-lm-imdb-promptstext10K<n<100K0 likes14k downloads4mo agoHugging Face05timaeus /rl-lm-toxicity-promptstext10K<n<100K0 likes8.9k downloads4mo agoHugging Face06nlile /NuminaMath-1.5-RL-Verifiable Dataset Card for NuminaMath-1.5-RL-Verifiable Dataset Summary NuminaMath-1.5-RL-Verifiable is a curated subset of the NuminaMath-1.5 dataset, specifically filtered to support reinforcement learning applications requiring verifiable outcomes. This collection consists of 131,063 math word problems from the original dataset that meet strict filtering criteria: all problems have definitive numerical answers, validated problem statements and solutions, and come from… See the full description on the dataset page: https://huggingface.co/datasets/nlile/NuminaMath-1.5-RL-Verifiable.texttext-generation100K<n<1M10 likes8.7k downloads1y agoHugging Face07lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909-completion Matched no-conftest RLVR study 20260909-completion Lossless research records, grouped by model and trajectory type. Only the listed configurations have published records. Canary diagnostics are excluded from study estimates; run status in provenance distinguishes retired diagnostics from active or completed training. Valid failures, refusals and truncations are retained. The train split name is a dataset-loader convention; record_type identifies whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.texttext-generation10K<n<100K1 likes7.2k downloads10d agoHugging Face08gavinlaw /rl-run-archive-2026 RL run archive 2026 Archived raw run artifacts (rollout trajectories, rendered frames, policy and optimizer checkpoints, configs, logs) from simulation reinforcement-learning experiments, published for long-term preservation and reproducibility. Layout mirrors the verified backup trees they were copied from: tilde/20260915-102000/ and taurus/20260915-085631/: batched tar archives. Every archive carries a per-file SHA-256 manifest inside it; the batch inventories (9998.json.gz… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/rl-run-archive-2026.tabularn<1K0 likes6.6k downloads2d agoHugging Face09RL-MIND /XHRBench XHRBench Ultra-High-Resolution Remote Sensing Understanding and Reasoning 🤗 Hugging Face · 🤖 ModelScope · 📄 Paper · 💻 Code English | 中文 📚 Introduction XHRBench evaluates fine-grained perception and complex reasoning in multimodal large language models using ultra-high-resolution remote-sensing imagery. This repository retains the name XHRBench and belongs to the same RSHR benchmark project as RSHR-Bench, with a… See the full description on the dataset page: https://huggingface.co/datasets/RL-MIND/XHRBench.imageimage-text-to-text1K<n<10K7 likes6.1k downloads6d agoHugging Face10AffineFoundation /rl-pythontext10K<n<100K2 likes5.8k downloads10mo agoHugging Face11PrimeIntellect /Reverse-Text-RL Reverse-Text-RL A small, scrappy RL dataset used in prime-rl's CI to debug RL training asking a model to reverse small sentences character-by-character. Follows the general format of PrimeIntellect/Reverse-Text-SFT The following script was used to generate the dataset. from datasets import Dataset, load_dataset dataset = load_dataset("willcb/R1-reverse-wikipedia-paragraphs-v1-1000", split="train") prompt = "Reverse the text character-by-character. Put your answer in… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Reverse-Text-RL.textquestion-answering1K<n<10K2 likes5.8k downloads1y agoHugging Face12SynthLabsAI /Big-Math-RL-Verifiedgated Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models Big-Math is the largest open-source dataset of high-quality mathematical problems, curated specifically for reinforcement learning (RL) training in language models. With over 250,000 rigorously filtered and verified problems, Big-Math bridges the gap between quality and quantity, establishing a robust foundation for advancing reasoning in LLMs. Request Early Access to Private… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/Big-Math-RL-Verified.textquestion-answering100K<n<1M243 likes5.4k downloads1y agoHugging Face13PrimeIntellect /Multi-SWE-RL-Verified Multi-SWE-RL-Verified Gold-patch-validated subset of PrimeIntellect/Multi-SWE-RL-Reupload (ByteDance's Multi-SWE-RL): 2,232 / 4,703 rows across C, Go, Java, JavaScript, Rust, and TypeScript that produce a clean reward signal end-to-end. Default dataset of the multiswe_v1 taskset. Changes vs upstream Starting from the 4,703-row re-upload: C++ dropped wholesale — 0/449 rows passed gold-patch validation in pass 1; the images are broken for scoring, not merely… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Multi-SWE-RL-Verified.tabulartext-generation1K<n<10K4 likes5.3k downloads3mo agoHugging Face14rl-rag /hle-gpt-oss-120b-no-python-260222 hle-gpt-oss-120b-no-python-260222 Deep research agent evaluation on rl-rag/hle_text_only (test split). Results Metric Value pass@4 47.9% avg@4 26.6% Trajectory accuracy 26.6% (2292/8632) Questions 2158 Trajectories 8632 (4 per question) Avg tool calls 14.5 Full conversations ❌ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domains huggingface.co Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-no-python-260222.tabular1K<n<10K1 likes5.1k downloads7mo agoHugging Face15KodCode /KodCode-Light-RL-10K 🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning. 🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-Light-RL-10K.tabularquestion-answering10K<n<100K9 likes4.6k downloads1y agoHugging Face16inclusionAI /ZwZ-RL-VQA ZwZ-RL-VQA: Region-to-Image Distilled Training Data for Fine-Grained Perception This synthetic dataset is generated via Region-to-Image Distillation (R2I) for training multimodal large language models (MLLMs) on fine-grained perception tasks without test-time tool use. 📖 Overview The Zooming without Zooming (ZwZ) method transforms "zooming" from an inference-time tool into a training-time primitive: Zoom-in Synthesis: Strong teacher models (Qwen3-VL-235B, GLM-4.5V)… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/ZwZ-RL-VQA.text100K<n<1M17 likes4.5k downloads4mo agoHugging Face17Stage-jh-monitor /appworld-qwen35-4b-agent-rl-epoch3 appworld-qwen35-4b-agent-rl-epoch3 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.45859375 Action score: 0.475 Valid samples: 320/320 tabularn<1K0 likes3.7k downloads13d agoHugging Face18Stage-jh-monitor /appworld-qwen35-4b-agent-rl-epoch3-reeval1 appworld-qwen35-4b-agent-rl-epoch3-reeval1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.4578125 Action score: 0.4921875 Valid samples: 320/320 tabularn<1K0 likes3.7k downloads13d agoHugging Face19huggingface-projects /Deep-RL-Course-Certificationtabular1K<n<10K19 likes3.6k downloads3h agoHugging Face20nvidia /Nemotron-RL-agent-workplace_assistant Dataset Description: The Nemotron-RL-agent-workplace_assistant is a tool use - multi step agentic environment that tests the agent’s ability to execute tasks in a workplace setting. Workbench contains a sandbox environment with five databases, 26 tools, and 690 tasks. These tasks represent common business activities, such as sending emails, scheduling meetings, etc. This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-agent-workplace_assistant.text1K<n<10K31 likes3.3k downloads7mo agoHugging Face21Lego-X /Lego-RL-2699 SWE-Lego-RL-2699 2,699 executable, difficulty-filtered SWE tasks for agentic RL, shipped in two parallel views of the same instances: View Path What it is Official OpenSWE records openswe_official_2699/ The original upstream GAIR/OpenSWE rows for exactly these 2,699 instances Harbor RL environments openswe_harbor_2699/ The same instances converted into ready-to-run task directories (+ the training index) Both views cover the identical 2,699 instance_ids. The… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/Lego-RL-2699.texttext-generation1K<n<10K2 likes3.2k downloads28d agoHugging Face22PrimeIntellect /INTELLECT-3-RLtabular10K<n<100K8 likes3.2k downloads4mo agoHugging Face23danil-e /rlpinn-ablation-runs RLPINN ablation runs Логи и результаты запусков абляции DQN-стека RL-агента (PINNacle). Буферы для этих запусков лежат в danil-e/rlpinn-ablation-buffers. Структура runs/<pde>/<ablation>/<run_tag>/ params.json # гиперпараметры запуска metrics.jsonl # по строке на лог метрик (шаг + ~30 метрик агента) others.json log.txt # полный stdout/stderr запуска rl_model_snapshots/ # веса агента по шагам… See the full description on the dataset page: https://huggingface.co/datasets/danil-e/rlpinn-ablation-runs.image1K<n<10K2 likes2.9k downloads2h agoHugging Face24Zchu /REDSearcher_RL_1Ktextquestion-answering1K<n<10K5 likes2.8k downloads2mo agoHugging Face25SLoonker /RL-Claude-Creative-Writing-SFT RL-Claude-Creative-Writing-SFT Alpaca-format dataset. Columns: instruction, input, output from datasets import load_dataset ds = load_dataset("SLoonker/RL-Claude-Creative-Writing-SFT", split="train") textn<1K1 likes2.8k downloads7mo agoHugging Face26IFM /guru-RL-92k Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Dataset Description Guru is a curated six-domain dataset for training large language models (LLM) for complex reasoning with reinforcement learning (RL). The dataset contains 91.9K high-quality samples spanning six diverse reasoning-intensive domains, processed through a comprehensive five-stage curation pipeline to ensure both domain diversity and reward verifiability.… See the full description on the dataset page: https://huggingface.co/datasets/IFM/guru-RL-92k.tabular10K<n<100K48 likes2.2k downloads1y agoHugging Face27allenai /RLVR-IFeval IF Data - RLVR Formatted This dataset contains instruction following data formatted for use with open-instruct - specifically reinforcement learning with verifiable rewards. Prompts with verifiable constraints generated by sampling from the Tulu 2 SFT mixture and randomly adding constraints from IFEval. Part of the Tulu 3 release, for which you can see models here and datasets here. Dataset Structure Each example in the dataset contains the standard instruction-tuning… See the full description on the dataset page: https://huggingface.co/datasets/allenai/RLVR-IFeval.text10K<n<100K36 likes2.2k downloads2y agoHugging Face28nvidia /Nemotron-RL-Agentic-Function-Calling-Pivot-v1 Dataset Description: This is a RL dataset for general function-calling by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Function-Calling-Pivot-v1.text1K<n<10K14 likes2k downloads7mo agoHugging Face29openbmb /RLAIF-V-Dataset Dataset Card for RLAIF-V-Dataset This dataset was introduced in RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness. GitHub This dataset was also used in MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe News: [2025.09.18] 🎉 Our data is used in the powerful MiniCPM-V 4.5 model, which represents a state-of-the-art end-side MLLM achieving GPT-4o level performance! [2025.03.01] 🎉 RLAIF-V is accepted by CVPR… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/RLAIF-V-Dataset.imageimage-text-to-text10K<n<100K219 likes1.9k downloads11mo agoHugging Face30nvidia /Nemotron-RL-Agentic-Terminal-Pivot-v1 Dataset Description The Nemotron-RL-Agentic-Terminal-Pivot-v1 dataset provides training samples for reinforcement learning of command-line ("terminal use") LLM agents with the terminus_judge environment in NeMo Gym. Each record is a single agent decision point extracted from a successful agent trajectory on a terminal task: responses_create_params.input — the prompt: the task instruction plus the terminal interaction history (prior agent actions and terminal outputs) up to the… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Terminal-Pivot-v1.texttext-generation10K<n<100K31 likes1.8k downloads25d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.