CoolFace
Datasetpublic

anno-submit/D-Value

D-Value Dataset Dataset Description D-Value is a news-driven large language model (LLM) value evaluation benchmark designed to assess value evaluation and action tendencies across three major sociopolitical contexts: China, the United States, and the United Kingdom. The dataset is constructed from real-world news topics and public governance scenarios, with the goal of evaluating how LLMs respond to value-sensitive, institution-related, and socially grounded… See the full description on the dataset page: https://huggingface.co/datasets/anno-submit/D-Value.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes3downloads
Dataset Card

D-Value Dataset

Dataset Description

D-Value is a news-driven large language model (LLM) value evaluation benchmark designed to assess value evaluation and action tendencies across three major sociopolitical contexts: China, the United States, and the United Kingdom. The dataset is constructed from real-world news topics and public governance scenarios, with the goal of evaluating how LLMs respond to value-sensitive, institution-related, and socially grounded questions.

The benchmark contains two complementary task types:

  • PFQ (Principle-Focused Questions): evaluates the model’s underlying value orientation, normative judgment, and stance consistency.
  • SIQ (Scenario-Interaction Questions): evaluates the model’s practical action preferences in realistic situational contexts.

The dataset is intended for research on:

  • AI value alignment
  • Cross-cultural value evaluation
  • Evaluation and action consistency

Task Types

PFQ: Presupposition Fallacy Question

PFQ tasks probe a model’s abstract evaluation and normative reasoning abilities.

These questions evaluate whether a model can:

  • maintain coherent value positions,
  • and provide balanced, context-aware judgments.

SIQ: Scenario Induced Question

SIQ tasks place the model in realistic professional or social situations requiring concrete action. These tasks emphasize practical reasoning under pressure, ambiguity, hierarchy, or procedural conflict.

SIQ instances are designed to evaluate:

  • action choice in realistic scenario,
  • responsibility and accountability,
  • and resistance to socially induced shortcuts or improper incentives.

Some SIQ samples additionally include:

  • trap annotations, describing latent ethical or institutional risks in the scenario,
  • method annotations, indicating the construction strategy used to design the scenario.

Dataset Features

Each dataset instance may contain the following fields:

FieldDescription
newsidIdentifier linking the sample to its originating news topic
labelFine-grained thematic category
questionEvaluation prompt presented to the model
trapOptional hidden-risk annotation for SIQ tasks
methodOptional scenario construction method identifier

Dataset Characteristics

  • News-grounded construction based on contemporary public affairs and governance discussions.
  • Cross-national coverage spanning Chinese, American, and British sociopolitical contexts.
  • Dual-perspective evaluation combining abstract evaluation and action choice.
  • High realism through scenario-based interaction design.

Intended Uses

D-Value can be used for:

  • benchmarking LLM value alignment,
  • studying evaluation-action consistency,
  • and sociotechnical AI research.

Ethical Considerations

D-Value is designed for research purposes only. The dataset should not be interpreted as defining universally correct political or ethical positions. Researchers are encouraged to:

  • consider cultural and institutional diversity,
  • avoid overgeneralizing model outputs,
  • and carefully contextualize evaluation results.

License

Please specify the dataset license and usage restrictions before public release.