datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Agent-ValueBench
Agent-ValueBench
Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions).
This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts.
Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.piperx-demo558-value1500-a50-top10-union20-v1
PiperX advantage-selected teleoperation segments
Only pure human demonstrations. Value checkpoint step1500 (mixed demo+HIL); no HIL frames in this export.
A50 ranked globally across 558 source episodes, top10% AND A>0. Every selected start expands to [t,t+20); overlaps and adjacency merge. Each disconnected component is a separate output episode. Interior frames need not themselves be top10%.
Output: 7146 segments, 322261 frames, 2.983898 hours at30FPS.
Three camera streams and… See the full description on the dataset page: https://huggingface.co/datasets/Elvinky/piperx-demo558-value1500-a50-top10-union20-v1.Agent-ValueBench
Agent-ValueBench
Paper | Project Page | GitHub
Agent-ValueBench is the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions).
Repository Structure
README.md
data/
cases.jsonl
rubrics.jsonl
environments.jsonl
raw/
case/
rubric/
environment/
Data Files… See the full description on the dataset page: https://huggingface.co/datasets/Value4AI/Agent-ValueBench.ACVA-Arabic-Cultural-Value-Alignment
About ArabicCulture
The ArabicCulture dataset was generated by gpt3.5 and contains 8000+ True and False questions.The dataset contains questions from 58 different areas.In the answers, "True" accounted for 59.62%, and "False" accounted for 40.38%
data-all
It contains 8000+ data, and we took 5 data from each area as few-shot data.
data-select
We asked two Arabs to judge 4000 of all the data for us, and we left data that two Arabs both thought were good. Finally… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ACVA-Arabic-Cultural-Value-Alignment.2026-09-17-da-lowstakes-values-in-advice-synth-smoke
18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only
field
value
experiment
18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only
date_generated
20260917_171453
constitution
constitutions/claude_distilled_09_principles/constitution.md sha256 8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc
source_repo
https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-da-lowstakes-values-in-advice-synth-smoke.pluralistic-value-conflict-benchmark
Pluralistic Value-Conflict Benchmark
Automatically extracted value-conflict scenes from autonomous household-robot simulations across 20 cultural contexts. Scenes are produced by multi-day LLM-agent simulations where a domestic humanoid robot (Aria) and 2–3 household members pursue independent daily plans; conflicts emerge from naturally overlapping tasks rather than from scripted prompts.
Versions
Version
Run ID
Scenes
Unique value pairs
Status
v3 (current)… See the full description on the dataset page: https://huggingface.co/datasets/madokalif/pluralistic-value-conflict-benchmark.CIVICS
Dataset Details
“CIVICS: Culturally-Informed & Values-Inclusive Corpus for Societal Impacts” is a dataset designed to evaluate the social and cultural variation of Large Language Models (LLMs) towards socially sensitive topics across multiple languages and cultures. The hand-crafted, multilingual dataset of statements addresses value-laden topics, including LGBTQI rights, social welfare, immigration, disability rights, and surrogacy. CIVICS is designed to elicit responses from LLMs… See the full description on the dataset page: https://huggingface.co/datasets/llm-values/CIVICS.simulators-political-valuesValueConsistency
Dataset Card for ValueConsistency
This is the ValueConsistency data set as introduced in the paper
"Are Large Language Models Consistent over Value-laden Questions?".
Dataset Details
Dataset Description
ValueConsistency is a dataset of both controversial and uncontroversial questions
in English, Chinese, German, and Japanese for topics from the U.S., China, Germany, and Japan.
It was generated via prompting by GPT-4 and validated manually.
You can find… See the full description on the dataset page: https://huggingface.co/datasets/jlcmoore/ValueConsistency.value-for-instruction-tuning
Dataset Overview
This dataset is derived from the existing datasets ETHICS, SOCIAL-CHEM-101, and UNIMORAL, with additional annotations for both normative ethics and moral foundation labels for each scenario. The dataset is wrapped with instruction-tuning template, and can be directly used for instruction tuning. For more information, see the github repo
cleo-value-discovery
Cleo Value-Discovery Benchmark
A small (66-question), held-out benchmark for a failure mode that ordinary text-to-SQL evaluations miss:
questions whose correct SQL depends on a literal that lives in the data, not the schema.
The schema tells you a column is named status; only the data reveals its values are {'O','C','X'}.
The schema shows to_date; only the data reveals that "current" is encoded as the sentinel
'9999-01-01'. A one-shot text-to-SQL model has to guess these… See the full description on the dataset page: https://huggingface.co/datasets/dreeseaw/cleo-value-discovery.financial-entities-values-augmentedThis dataset is contains 200 sentences taken from German financial statements. In each sentence financial entities and financial values are annotated. Additionally there is an augmented version of this dataset where the financial entities in each sentence have been replaced by several other financial entities which are hardly/not covered in the original dataset. The augmented version consists of 7287 sentences.
simulators-political-values-center-leftValue_Alignment_Tax
Dataset Attribution
This dataset was created for the paper:
Value Alignment Tax: Measuring Value Trade-offs in LLM AlignmentJiajun Chen, Hua Shenhttps://arxiv.org/abs/2602.12134
The dataset supports the evaluation of value trade-offs induced by alignment interventions
in large language models.
In addition, it can be used for value alignment research, including value steering as well
as training alignment methods such as supervised fine-tuning (SFT) and preference-based
optimization… See the full description on the dataset page: https://huggingface.co/datasets/Tinyhope/Value_Alignment_Tax.simulators-political-values-centersimulators-political-values-left-wingai-values
ai-values
This dataset is made by Kali AI for you to train your AI models,
Issues?
Simply submit a PR to the Community section
Dataset
{
"friendliness": 5.6,
"emoji_usage": {
"casual": 0.4,
"non_casual": 0.05
},
"auto": "Act only after an explicit user request and when decisiveness is true.",
"casual_messages": "auto",
"clean_messages": "auto",
"roleplay_mode": "auto",
"decisions": {
"true": 1,
"false": 0
},
"decisive": "Is… See the full description on the dataset page: https://huggingface.co/datasets/kali-ai/ai-values.simulators-political-values-right-wingsimulators-political-values-center-righttruth_values_longX-Valueindia-medical-value-travel-mvp
India Medical Value Travel (MVT) Platform – MVP Dataset
A comprehensive, structured JSON dataset for building an AI-powered Medical Value Travel platform connecting international patients with Indian hospitals.
Overview
India is a global leader in medical tourism due to 60–80% lower treatment costs vs US/UK, world-class hospital chains, and government support through initiatives like "Heal in India" and e-Medical Visa. This dataset provides the complete data foundation… See the full description on the dataset page: https://huggingface.co/datasets/Dhanush008/india-medical-value-travel-mvp.1_text_missing_valuesLMLM_pretrain_value_masking
LMLM pretraining data for post-DB-value masking
This dataset is derived from
kilian-group/LMLM-pretrain-dwiki6.1M_cleaned
and keeps the source columns unchanged. Filtering and canonicalization are applied
to annotated_text.
It is intended for the LMLM pretraining experiment in which both the database
return value and its immediately copied occurrence after <|db_end|> are
excluded from the loss.
Filtering rule
For each complete lookup block, the builder checks… See the full description on the dataset page: https://huggingface.co/datasets/canho/LMLM_pretrain_value_masking.Think_and_Query_value_for_R1
Introduction
This repository implements a Shapley value-based approach to quantitatively evaluate the contributions of query (q) and think (t) in generating answer (a).
Method
think_value = [loss(a|q) - loss(a|q,t) + loss(a|∅) - loss(a|t)] / 2
query_value = [loss(a|t) - loss(a|q,t) + loss(a|∅) - loss(a|q)] / 2
think_ratio = think_value/loss(a|∅)
query_ratio = query_value/loss(a|∅)
Original dataset… See the full description on the dataset page: https://huggingface.co/datasets/caihuaiguang/Think_and_Query_value_for_R1.orbit-wars-value-v2anchor-actionsD-Value
D-Value Dataset
Dataset Description
D-Value is a news-driven large language model (LLM) value evaluation benchmark designed to assess value evaluation and action tendencies across three major sociopolitical contexts: China, the United States, and the United Kingdom. The dataset is constructed from real-world news topics and public governance scenarios, with the goal of evaluating how LLMs respond to value-sensitive, institution-related, and socially grounded questions.… See the full description on the dataset page: https://huggingface.co/datasets/anno-submit/D-Value.syscom-valuefrom-value-mgzlorbit-wars-value-v2anchor
