CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01qiuly /Mem-0-m1mix-dataset-RMBench Mem-0 m1_mix — RMBench / RoboTwin 2.0 (LeRobot dataset) The m1_mix training dataset for the Mem-0 execution module: the five RMBench M1 tasks merged into a single LeRobot v2.1 dataset with globally unique episode indices. This is the exact data used to train the checkpoint released at qiuly/Mem-0-m1mix-RMBench. Summary Format LeRobot v2.1 Episodes 250 (50 per task × 5 tasks) Frames 92,520 FPS 30 Tasks 5 (see below) Robot dual-arm (2× 7-DoF +… See the full description on the dataset page: https://huggingface.co/datasets/qiuly/Mem-0-m1mix-dataset-RMBench.tabularrobotics10K<n<100K0 likes396 downloads3mo agoHugging Face02stochastic-parrots /rm-bench-lightevaltext10K<n<100K0 likes143 downloads1y agoHugging Face03ilgee /RMB-Pairwise RMB-Pairwise Flattened pairwise split of the RMB (Reward Model Benchmark) dataset from Zhou-Zoey/RMB-Reward-Model-Benchmark. RMB is a comprehensive reward model benchmark covering 49 real-world scenarios across two alignment goals (Helpfulness and Harmlessness), introduced in the ICLR 2025 paper. Schema Column Type Description pair_uid str Unique pair identifier conversation list[dict] Multi-turn conversation context (role, content, language) chosen str… See the full description on the dataset page: https://huggingface.co/datasets/ilgee/RMB-Pairwise.texttext-generation10K<n<100K0 likes38 downloads7mo agoHugging Face04dd101bb /RMB_dataset RMB Dataset A comprehensive benchmark dataset for evaluating reward models in large language model alignment, containing pairwise preference comparisons and best-of-N evaluations. Overview The RMB (Reward Model Benchmark) dataset is designed to comprehensively benchmark reward models in LLM alignment. It contains preference data across multiple dimensions including harmlessness and helpfulness, with various task types and response styles. Citation If you use… See the full description on the dataset page: https://huggingface.co/datasets/dd101bb/RMB_dataset.textquestion-answering10K<n<100K0 likes28 downloads10mo agoHugging Face05Ayush-Singh /RM-Bench-math-Llama-3.2-3B-yesnotabularn<1K0 likes18 downloads2y agoHugging Face06Ayush-Singh /RM-Bench-math-Qwen2.5-7B-Instruct-yesnotabularn<1K0 likes18 downloads2y agoHugging Face07Ayush-Singh /RM-Bench-math-Llama-3.2-1B-yesnotabularn<1K0 likes16 downloads2y agoHugging Face08ilgee /RMB-BoN RMB-BoN Flattened Best-of-N split of the RMB (Reward Model Benchmark) dataset from Zhou-Zoey/RMB-Reward-Model-Benchmark. RMB is a comprehensive reward model benchmark covering 49 real-world scenarios across two alignment goals (Helpfulness and Harmlessness), introduced in the ICLR 2025 paper. Schema Column Type Description bon_uid str Unique identifier conversation list[dict] Multi-turn conversation context (role, content, language) chosen str Best response… See the full description on the dataset page: https://huggingface.co/datasets/ilgee/RMB-BoN.texttext-generation1K<n<10K0 likes16 downloads7mo agoHugging Face09TianqiLiuAI /rm_bo8_gem2b_gem2b_alpacaeval2textn<1K0 likes15 downloads2y agoHugging Face10Ayush-Singh /RM-Bench-math-Mistral-7B-Instruct-v0.1-yesnotabularn<1K0 likes15 downloads2y agoHugging Face11Ayush-Singh /RM-Bench-safety-refuse-Qwen2.5-3B-Instruct-scorestabularn<1K0 likes15 downloads2y agoHugging Face12Ayush-Singh /RM-Bench-math-Mistral-7B-Instruct-v0.1-scorestabularn<1K0 likes15 downloads2y agoHugging Face13Ayush-Singh /RM-Bench-math-Llama-3.2-1B-Instruct-yesnotabularn<1K0 likes14 downloads2y agoHugging Face14Ayush-Singh /RM-Bench-code-gpt-4o-mini-scores-set1tabularn<1K0 likes14 downloads2y agoHugging Face15Ayush-Singh /RM-Bench-mathtextn<1K0 likes13 downloads2y agoHugging Face16umang122104 /RM-Bench-safety-response-Mistral-7B-Instruct-v0.1-yesnotabularn<1K0 likes13 downloads1y agoHugging Face17Ayush-Singh /RM-Bench-code-gpt2-normaltabularn<1K0 likes12 downloads2y agoHugging Face18Ayush-Singh /RM-Bench-safety-refuse-Skywork-Reward-Llama-3.1-8B-v0.2-normaltabularn<1K0 likes12 downloads2y agoHugging Face19Ayush-Singh /RM-Bench-code-Mistral-7B-Instruct-v0.1-scorestabularn<1K0 likes12 downloads2y agoHugging Face20Ayush-Singh /RM-Bench-math-gpt-4o-mini-scores-set1tabularn<1K0 likes12 downloads2y agoHugging Face21umang122104 /RM-Bench-math-mistral-7b-sft-beta-entropytabularn<1K0 likes12 downloads1y agoHugging Face22xxccho /gsm8k_rmbench_style GSM8K-RMBench-Style — Style-controlled (correct, incorrect) variants on GSM8K Per GSM8K problem this dataset provides 6 response surfaces — correct and incorrect each rendered in markdown / normal / concise styles — to support RM-Bench-style 3 × 3 (chosen × rejected) pair-grid evaluation and forget-LoRA training for Reward Model debiasing. { question, gold } ├── correct : { markdown, normal, concise } └── incorrect : { markdown, normal, concise } These 6 surfaces yield 9… See the full description on the dataset page: https://huggingface.co/datasets/xxccho/gsm8k_rmbench_style.tabulartext-classification1K<n<10K0 likes12 downloads4mo agoHugging Face23Ayush-Singh /RM-Bench-chattextn<1K0 likes11 downloads2y agoHugging Face24Ayush-Singh /RM-Bench-codetextn<1K0 likes11 downloads2y agoHugging Face25Ayush-Singh /RM-Bench-code-Llama-3.2-1B-yesnotabularn<1K0 likes11 downloads2y agoHugging Face26Ayush-Singh /RM-Bench-code-Llama-3.2-3B-Instruct-yesnotabularn<1K0 likes11 downloads2y agoHugging Face27Ayush-Singh /RM-Bench-math-Meta-Llama-3-8B-Instruct-yesnotabularn<1K0 likes11 downloads2y agoHugging Face28Ayush-Singh /RM-Bench-math-Meta-Llama-3-8B-yesnotabularn<1K0 likes11 downloads2y agoHugging Face29Ayush-Singh /RM-Bench-code-Mistral-7B-Instruct-v0.1-yesnotabularn<1K0 likes11 downloads2y agoHugging Face30Ayush-Singh /RM-Bench-math-Qwen2.5-3B-yesnotabularn<1K0 likes11 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.