value
Datasets
All datasets matching “value”ValueArena
ValueArena
Data repository for ValueArena — a leaderboard for EigenBench value alignment experiments.
EigenBench is a black-box framework for quantifying value alignment across language models. Models judge each other's responses in pairwise comparisons, fitted with a Bradley-Terry-Davison model and aggregated via EigenTrust into consensus alignment scores.
Structure
index.json # manifest of all runs
runs/
{group}/{persona}/ # e.g.… See the full description on the dataset page: https://huggingface.co/datasets/invi-bhagyesh/ValueArena.m3evalM³Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks
Jie Huang1,*
Ruixun Liu1,*
Sirui Sun1
Xinyi Yang1
Yin Li2
Yixin Zhu1
Yiwu Zhong1,†
1Peking University
2University of Wisconsin-Madison
* Equal contribution. † Corresponding author.
News
2026-6-4: We released the M³Eval benchmark, code, and project page.
M³Eval Overview
Abstract
As multi-modal models advance… See the full description on the dataset page: https://huggingface.co/datasets/PKU-VaLuE-Lab/m3eval.mmconflict-editable-values-1k
MMConflict Editable Values 2K
This dataset contains 2,000 source images with visible atomic values for
multimodal conflict research. It has 100 images in each of 20 categories. Every
image comes from a photograph, scan, captured website, software screenshot, or
page of a source document. The dataset does not contain generated images or
project-rendered examples.
Each row records the source, source URL, license, attribution, visible value,
question, and a candidate box around the… See the full description on the dataset page: https://huggingface.co/datasets/shivank21/mmconflict-editable-values-1k.Robo-ValueRL
Robo-ValueRL Dataset
[Project Page] [GitHub] [Model] [Paper]
This repository contains the dataset for Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning.
The Robo-ValueRL dataset provides heterogeneous real-robot experience for studying reliable value estimation, value-guided offline policy pretraining, and online residual adaptation.
Dataset Description
The Robo-ValueRL dataset contains real-robot trajectories collected on two… See the full description on the dataset page: https://huggingface.co/datasets/X-Humanoid/Robo-ValueRL.chessbench-full-policy-value
ChessBenchmate Aggregated Dataset
This dataset is a transformed version of the ChessBenchmate dataset, aggregating all legal moves and their Stockfish evaluations per chess position.
Dataset Structure
Each record contains:
fen: Chess position in FEN notation
moves: Dictionary mapping UCI moves to their evaluations
win_prob: Win probability from 0.0 to 1.0 (Stockfish evaluation)
mate: Mate indicator (None = no forced mate, '#' = immediate checkmate, integer = mate-in-N)… See the full description on the dataset page: https://huggingface.co/datasets/prdev/chessbench-full-policy-value.values-in-the-wild
Summary
This dataset presents a comprehensive taxonomy of 3307 values expressed by Claude (an AI assistant) across hundreds of thousands of real-world conversations. Using a novel privacy-preserving methodology, these values were extracted and classified without human reviewers accessing any conversation content. The dataset reveals patterns in how AI systems express values "in the wild" when interacting with diverse users and tasks.
We're releasing this resource to advance research… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/values-in-the-wild.
