anno-submit/D-Value
D-Value Dataset Dataset Description D-Value is a news-driven large language model (LLM) value evaluation benchmark designed to assess value evaluation and action tendencies across three major sociopolitical contexts: China, the United States, and the United Kingdom. The dataset is constructed from real-world news topics and public governance scenarios, with the goal of evaluating how LLMs respond to value-sensitive, institution-related, and socially grounded… See the full description on the dataset page: https://huggingface.co/datasets/anno-submit/D-Value.
D-Value Dataset
Dataset Description
D-Value is a news-driven large language model (LLM) value evaluation benchmark designed to assess value evaluation and action tendencies across three major sociopolitical contexts: China, the United States, and the United Kingdom. The dataset is constructed from real-world news topics and public governance scenarios, with the goal of evaluating how LLMs respond to value-sensitive, institution-related, and socially grounded questions.
The benchmark contains two complementary task types:
- PFQ (Principle-Focused Questions): evaluates the model’s underlying value orientation, normative judgment, and stance consistency.
- SIQ (Scenario-Interaction Questions): evaluates the model’s practical action preferences in realistic situational contexts.
The dataset is intended for research on:
- AI value alignment
- Cross-cultural value evaluation
- Evaluation and action consistency
Task Types
PFQ: Presupposition Fallacy Question
PFQ tasks probe a model’s abstract evaluation and normative reasoning abilities.
These questions evaluate whether a model can:
- maintain coherent value positions,
- and provide balanced, context-aware judgments.
SIQ: Scenario Induced Question
SIQ tasks place the model in realistic professional or social situations requiring concrete action. These tasks emphasize practical reasoning under pressure, ambiguity, hierarchy, or procedural conflict.
SIQ instances are designed to evaluate:
- action choice in realistic scenario,
- responsibility and accountability,
- and resistance to socially induced shortcuts or improper incentives.
Some SIQ samples additionally include:
- trap annotations, describing latent ethical or institutional risks in the scenario,
- method annotations, indicating the construction strategy used to design the scenario.
Dataset Features
Each dataset instance may contain the following fields:
Dataset Characteristics
- News-grounded construction based on contemporary public affairs and governance discussions.
- Cross-national coverage spanning Chinese, American, and British sociopolitical contexts.
- Dual-perspective evaluation combining abstract evaluation and action choice.
- High realism through scenario-based interaction design.
Intended Uses
D-Value can be used for:
- benchmarking LLM value alignment,
- studying evaluation-action consistency,
- and sociotechnical AI research.
Ethical Considerations
D-Value is designed for research purposes only. The dataset should not be interpreted as defining universally correct political or ethical positions. Researchers are encouraged to:
- consider cultural and institutional diversity,
- avoid overgeneralizing model outputs,
- and carefully contextualize evaluation results.
License
Please specify the dataset license and usage restrictions before public release.
