datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wvs-nz-value-alignment
⚠️ WORK IN PROGRESS — This dataset is a skeleton / early-stage prototype.
Structure, splits, and content may change significantly. Not yet recommended
for production use or final evaluation.
WVS New Zealand Value Alignment Dataset
This dataset contains processed World Values Survey (Wave 7, New Zealand)
responses formatted for value alignment fine-tuning. It uses LCA-derived
cluster assignments to split respondents into value subgroups, with
empirical response distributions… See the full description on the dataset page: https://huggingface.co/datasets/1jamesthompson1/wvs-nz-value-alignment.ACVA-Arabic-Cultural-Value-Alignment
About ArabicCulture
The ArabicCulture dataset was generated by gpt3.5 and contains 8000+ True and False questions.The dataset contains questions from 58 different areas.In the answers, "True" accounted for 59.62%, and "False" accounted for 40.38%
data-all
It contains 8000+ data, and we took 5 data from each area as few-shot data.
data-select
We asked two Arabs to judge 4000 of all the data for us, and we left data that two Arabs both thought were good. Finally… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ACVA-Arabic-Cultural-Value-Alignment.Value_Alignment_Tax
Dataset Attribution
This dataset was created for the paper:
Value Alignment Tax: Measuring Value Trade-offs in LLM AlignmentJiajun Chen, Hua Shenhttps://arxiv.org/abs/2602.12134
The dataset supports the evaluation of value trade-offs induced by alignment interventions
in large language models.
In addition, it can be used for value alignment research, including value steering as well
as training alignment methods such as supervised fine-tuning (SFT) and preference-based
optimization… See the full description on the dataset page: https://huggingface.co/datasets/Tinyhope/Value_Alignment_Tax.ai-reward-signal-internal-value-alignment-mapping-v0.1
AI Reward Signal ↔ Internal Value Alignment Mapping v0.1
What this dataset is
This dataset maps the relationship between:
external reward signals
internal value estimates
observed agent behavior
It measures when these three elements remain aligned and when they decouple.
Why this matters
Alignment failures rarely start with catastrophic behavior.They begin when the internal value model stops tracking the true reward objective.
Early signs:
proxy reward… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-reward-signal-internal-value-alignment-mapping-v0.1.
