CoolFace
20 results

value

invi-bhagyesh /ValueArena ValueArena Data repository for ValueArena — a leaderboard for EigenBench value alignment experiments. EigenBench is a black-box framework for quantifying value alignment across language models. Models judge each other's responses in pairwise comparisons, fitted with a Bradley-Terry-Davison model and aggregated via EigenTrust into consensus alignment scores. Structure index.json # manifest of all runs runs/ {group}/{persona}/ # e.g.… See the full description on the dataset page: https://huggingface.co/datasets/invi-bhagyesh/ValueArena.0 likes2.6k downloads4d agoHugging FacePKU-VaLuE-Lab /m3evalM³Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks Jie Huang1,*  Ruixun Liu1,*  Sirui Sun1  Xinyi Yang1  Yin Li2  Yixin Zhu1  Yiwu Zhong1,† 1Peking University  2University of Wisconsin-Madison * Equal contribution. † Corresponding author. News 2026-6-4: We released the M³Eval benchmark, code, and project page. M³Eval Overview Abstract As multi-modal models advance… See the full description on the dataset page: https://huggingface.co/datasets/PKU-VaLuE-Lab/m3eval.videovisual-question-answering1K<n<10K1 likes2.6k downloads4mo agoHugging Faceshivank21 /mmconflict-editable-values-1k MMConflict Editable Values 2K This dataset contains 2,000 source images with visible atomic values for multimodal conflict research. It has 100 images in each of 20 categories. Every image comes from a photograph, scan, captured website, software screenshot, or page of a source document. The dataset does not contain generated images or project-rendered examples. Each row records the source, source URL, license, attribution, visible value, question, and a candidate box around the… See the full description on the dataset page: https://huggingface.co/datasets/shivank21/mmconflict-editable-values-1k.imageimage-to-text1K<n<10K0 likes1.1k downloads21d agoHugging FaceX-Humanoid /Robo-ValueRL Robo-ValueRL Dataset [Project Page] [GitHub] [Model] [Paper] This repository contains the dataset for Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning. The Robo-ValueRL dataset provides heterogeneous real-robot experience for studying reliable value estimation, value-guided offline policy pretraining, and online residual adaptation. Dataset Description The Robo-ValueRL dataset contains real-robot trajectories collected on two… See the full description on the dataset page: https://huggingface.co/datasets/X-Humanoid/Robo-ValueRL.video10K<n<100K1 likes1k downloads3mo agoHugging Faceprdev /chessbench-full-policy-value ChessBenchmate Aggregated Dataset This dataset is a transformed version of the ChessBenchmate dataset, aggregating all legal moves and their Stockfish evaluations per chess position. Dataset Structure Each record contains: fen: Chess position in FEN notation moves: Dictionary mapping UCI moves to their evaluations win_prob: Win probability from 0.0 to 1.0 (Stockfish evaluation) mate: Mate indicator (None = no forced mate, '#' = immediate checkmate, integer = mate-in-N)… See the full description on the dataset page: https://huggingface.co/datasets/prdev/chessbench-full-policy-value.tabular-classification1B<n<10B0 likes879 downloads8mo agoHugging FaceAnthropic /values-in-the-wild Summary This dataset presents a comprehensive taxonomy of 3307 values expressed by Claude (an AI assistant) across hundreds of thousands of real-world conversations. Using a novel privacy-preserving methodology, these values were extracted and classified without human reviewers accessing any conversation content. The dataset reveals patterns in how AI systems express values "in the wild" when interacting with diverse users and tasks. We're releasing this resource to advance research… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/values-in-the-wild.tabular1K<n<10K155 likes764 downloads1y agoHugging Face