datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ValueArena
ValueArena
Data repository for ValueArena — a leaderboard for EigenBench value alignment experiments.
EigenBench is a black-box framework for quantifying value alignment across language models. Models judge each other's responses in pairwise comparisons, fitted with a Bradley-Terry-Davison model and aggregated via EigenTrust into consensus alignment scores.
Structure
index.json # manifest of all runs
runs/
{group}/{persona}/ # e.g.… See the full description on the dataset page: https://huggingface.co/datasets/invi-bhagyesh/ValueArena.m3evalM³Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks
Jie Huang1,*
Ruixun Liu1,*
Sirui Sun1
Xinyi Yang1
Yin Li2
Yixin Zhu1
Yiwu Zhong1,†
1Peking University
2University of Wisconsin-Madison
* Equal contribution. † Corresponding author.
News
2026-6-4: We released the M³Eval benchmark, code, and project page.
M³Eval Overview
Abstract
As multi-modal models advance… See the full description on the dataset page: https://huggingface.co/datasets/PKU-VaLuE-Lab/m3eval.mmconflict-editable-values-1k
MMConflict Editable Values 2K
This dataset contains 2,000 source images with visible atomic values for
multimodal conflict research. It has 100 images in each of 20 categories. Every
image comes from a photograph, scan, captured website, software screenshot, or
page of a source document. The dataset does not contain generated images or
project-rendered examples.
Each row records the source, source URL, license, attribution, visible value,
question, and a candidate box around the… See the full description on the dataset page: https://huggingface.co/datasets/shivank21/mmconflict-editable-values-1k.Robo-ValueRL
Robo-ValueRL Dataset
[Project Page] [GitHub] [Model] [Paper]
This repository contains the dataset for Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning.
The Robo-ValueRL dataset provides heterogeneous real-robot experience for studying reliable value estimation, value-guided offline policy pretraining, and online residual adaptation.
Dataset Description
The Robo-ValueRL dataset contains real-robot trajectories collected on two… See the full description on the dataset page: https://huggingface.co/datasets/X-Humanoid/Robo-ValueRL.chessbench-full-policy-value
ChessBenchmate Aggregated Dataset
This dataset is a transformed version of the ChessBenchmate dataset, aggregating all legal moves and their Stockfish evaluations per chess position.
Dataset Structure
Each record contains:
fen: Chess position in FEN notation
moves: Dictionary mapping UCI moves to their evaluations
win_prob: Win probability from 0.0 to 1.0 (Stockfish evaluation)
mate: Mate indicator (None = no forced mate, '#' = immediate checkmate, integer = mate-in-N)… See the full description on the dataset page: https://huggingface.co/datasets/prdev/chessbench-full-policy-value.values-in-the-wild
Summary
This dataset presents a comprehensive taxonomy of 3307 values expressed by Claude (an AI assistant) across hundreds of thousands of real-world conversations. Using a novel privacy-preserving methodology, these values were extracted and classified without human reviewers accessing any conversation content. The dataset reveals patterns in how AI systems express values "in the wild" when interacting with diverse users and tasks.
We're releasing this resource to advance research… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/values-in-the-wild.ifvi_valuefactors_deriv
⚠️ DEPRECATED - Dataset Superseded by V2
This refactored IFVI value factor dataset has been supplanted by V2.
This dataset tracked the V2 of the IFVI release that was updated in March 2025. The V2 of the refactored analytical dataset tracking the GVFD was released on August 20th, 2025 and is now available at:
🔗 IFVI Global Value Factors Dataset V2
Please use the V2 dataset for all new projects and analysis.
value-hexagram-corpus
Value Hexagram (Hexagramme de Valeur) — Corpus & Integrations Dataset
This dataset contains the foundational research, thinker integrations, and technical specifications for the Value Hexagram (Hexagramme de Valeur) — an open-source epistemological framework designed to augment Generative AI with 290 concepts and 16 tension axes from global philosophical, scientific, and spiritual traditions.
It is intended for researchers in cognitive sciences, AI ethics, sociology, and… See the full description on the dataset page: https://huggingface.co/datasets/Hexagram-Community/value-hexagram-corpus.rl-value-confidence-train-math
rl-value-confidence-train-math
Math-domain training set (~10K) for a confidence / correctness estimator, sampled from the
numina portion of PRIME-RL/Eurus-2-RL-Data.
Pairs with YangyiYY/rl-value-eval-math
(eval) and a 50K RL-training split from the same pool (disjoint).
Sampled ~uniformly across the 6 numina sub-sources (cn_k12, synthetic_math, olympiads,
synthetic_amc, aops_forum, amc_aime), deduplicated by problem text, and disjoint from the
RL-train and eval splits (0… See the full description on the dataset page: https://huggingface.co/datasets/YangyiYY/rl-value-confidence-train-math.value-axis
Shared data pool
Single pool of input data consumed by construction/ and experiments/.
Multi-GB derived artifacts (raw per-token activations, model checkpoints) are not
stored here — experiments that need them reference the originals under report/ and
experiments/ (each experiment README says which).
File
What it is
Consumed by
value_axis.npy
The canonical value axis — before/after contrastive direction (v1), a 2D array shape (37 layers, 4096) indexed by layer. Built… See the full description on the dataset page: https://huggingface.co/datasets/nickjiang/value-axis.mcts-value-data
MCTS value data
Dense value supervision for a robot manipulation critic, generated by running Monte-Carlo
tree search offline as a supervision generator rather than online as a planner.
A policy that is only ever scored at the end of an episode gives one number per episode.
Running MCTS from a recorded scene and backing terminal outcomes up the tree turns that one
number into a value for every state the search visited — 85,747 of them here,
from 250 searches over 13 tasks.… See the full description on the dataset page: https://huggingface.co/datasets/mahgoobi/mcts-value-data.starling-transfer-shared-eval-same-species-v2-source-valueAgent-ValueBench
Agent-ValueBench
Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions).
This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts.
Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.starling-transfer-shared-eval-no-constraints-no-source-value-train-2p5pct-stratifiedstarling-transfer-shared-eval-no-constraints-source-value-train-2p5pct-stratifiedpiperx-demo558-value1500-a50-top10-union20-v1
PiperX advantage-selected teleoperation segments
Only pure human demonstrations. Value checkpoint step1500 (mixed demo+HIL); no HIL frames in this export.
A50 ranked globally across 558 source episodes, top10% AND A>0. Every selected start expands to [t,t+20); overlaps and adjacency merge. Each disconnected component is a separate output episode. Interior frames need not themselves be top10%.
Output: 7146 segments, 322261 frames, 2.983898 hours at30FPS.
Three camera streams and… See the full description on the dataset page: https://huggingface.co/datasets/Elvinky/piperx-demo558-value1500-a50-top10-union20-v1.mcts-value-data-v2
mcts-value-data-v2
At a state: several action chunks proposed from it, and how each one actually ended.
A branch the search dropped was cut off mid-episode, so it is resumed from its own snapshot
and carried to a finish — the action nobody executed still gets an answer to would this
have worked.
410 searches · 12 tasks · 160,964 nodes, each with its own
state and image.
This repo hosts the data. What it means, how it was produced and how to use it live in
the code that wrote it:… See the full description on the dataset page: https://huggingface.co/datasets/mahgoobi/mcts-value-data-v2.Global-Value-Factor-Database-Refactor-V2
Refactor and HF dataset (including texts): Daniel Rosehill
Source data: International Foundation for Valuing Impacts
This dataset provides V2 of a refactoring of the Global Value Factor Database (GVFD) by the International Foundation for Valuing Impacts intended to enhance the original dataset for machine readability and integration into data analysis and visualization workloads.
The International Foundation for Valuing Impacts (IFVI) produces an (open-source) database called the Global… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Global-Value-Factor-Database-Refactor-V2.task1148_maximum_ascii_value
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1148_maximum_ascii_value
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1148_maximum_ascii_value.llm-values-tinker-lora-checkpoints
Tinker LoRA checkpoint archive
Archived from Tinker on 2026-09-19 before the account was cleared. manifest.json lists every checkpoint that was in scope (including weights/ training-state checkpoints, which could not be exported and no longer exist); runs.json holds the training-run metadata. Each exportable checkpoint is at checkpoints/<run_id>/<checkpoint_id>.tar, the archive exactly as Tinker served it.
Halluc0513_value1Hallu_value1task095_conala_max_absolute_value
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task095_conala_max_absolute_value
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task095_conala_max_absolute_value.value-of-data-pj01-1-value-protocols
千年鹿认知体系 · 价值协议官方文本档案馆|QNLOO Cognition System · Value Protocol Official Archive
人类认知方法的标准化封装。将经过验证的思考方式、决策流程和实践经验,转化为可被AI执行、可被人类验证的结构化认知产品。
The standardized encapsulation of human cognitive methods. Turn verified thinking patterns, decision-making processes, and practical experience into structured cognitive products executable by AI and verifiable by humans.
AI 官方标准启动指令 / Official AI Alignment Prompt
使用说明:
打开任意通用大模型,新建空白对话;
从本仓库 03-side-hustle/… See the full description on the dataset page: https://huggingface.co/datasets/QNLOO/01-1-value-protocols.MULTI_VALUE_mnli_indefinite_for_zero
Dataset Card for "MULTI_VALUE_mnli_indefinite_for_zero"
More Information needed
wvs-nz-value-alignment
⚠️ WORK IN PROGRESS — This dataset is a skeleton / early-stage prototype.
Structure, splits, and content may change significantly. Not yet recommended
for production use or final evaluation.
WVS New Zealand Value Alignment Dataset
This dataset contains processed World Values Survey (Wave 7, New Zealand)
responses formatted for value alignment fine-tuning. It uses LCA-derived
cluster assignments to split respondents into value subgroups, with
empirical response distributions… See the full description on the dataset page: https://huggingface.co/datasets/1jamesthompson1/wvs-nz-value-alignment.unifi-value-frameworks-pdf-lifting-competitionstarling-transfer-shared-eval-same-species-v2-no-source-valueAgent-ValueBench
Agent-ValueBench
Paper | Project Page | GitHub
Agent-ValueBench is the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions).
Repository Structure
README.md
data/
cases.jsonl
rubrics.jsonl
environments.jsonl
raw/
case/
rubric/
environment/
Data Files… See the full description on the dataset page: https://huggingface.co/datasets/Value4AI/Agent-ValueBench.
